The impending realization of cryptanalytically relevant quantum computers (CRQCs) threatens the foundational security architecture of global digital communications. Shor's algorithm provides a polynomial-time mechanism to solve prime factorization and discrete logarithms, rendering classical public-key cryptography—specifically RSA, Elliptic Curve Diffie-Hellman (ECDH), and ECDSA—completely vulnerable. Furthermore, adversarial "Store Now, Decrypt Later" (SNDL) espionage campaigns are actively archiving encrypted Transport Layer Security (TLS) network traffic to decrypt it retrospectively once fault-tolerant quantum hardware matures. In response to this urgent existential risk, the National Institute of Standards and Technology (NIST) finalized its principal post-quantum cryptographic standards in August 2024: FIPS 203 (ML-KEM / Module-Lattice Key Encapsulation Mechanism) and FIPS 204 (ML-DSA / Module-Lattice Digital Signature Algorithm). However, transitioning production internet services to post-quantum TLS 1.3 introduces severe engineering bottlenecks governed by polynomial arithmetic overhead over high-degree rings R_q = Z_q[X]/(X^256 + 1), substantial public key and signature sizes, L1/L2 cache thrashing, and wide-area network (WAN) packet fragmentation across standard Ethernet Maximum Transmission Units (MTUs). To resolve these architectural uncertainties, this paper presents a comprehensive, empirical microarchitectural and network benchmarking evaluation of ML-KEM-768 and ML-DSA-65 integrated within TLS 1.3 handshakes across three heterogeneous processor instruction set architectures (ISAs): Intel x86-64 Xeon (AVX-512), ARMv8/v9 Neoverse V2 (NEON), and SiFive RISC-V (RVV 1.0). We systematically profile SIMD vectorization of the Number Theoretic Transform (NTT), Montgomery/Barrett modular reductions, constant-time side-channel resistance under Welch's t-test (DudeCT), and network-layer TCP packet fragmentation under varying round-trip times (RTT) and packet loss rates. Our empirical findings establish that: • SIMD Microarchitectural Acceleration: Dedicated vectorization yields a 6.2x throughput speedup on x86 AVX-512, a 4.1x speedup on ARM NEON, and a 3.6x speedup on RISC-V Vector compared to unvectorized C reference implementations, reducing ML-KEM-768 key encapsulation to 41,200 CPU cycles. • Production TLS 1.3 Handshake Feasibility: The standard hybrid key exchange X25519MLKEM768 adds merely 0.38 ms of cryptographic compute latency on modern servers, operating well within real-world web application performance budgets. • Network Packet Fragmentation & Tail Latency: The large signature footprint of ML-DSA-65 (3,293 bytes) forces the ServerHello and certificate chain across multiple Ethernet MTU boundaries (3 TCP segments). Under 1% WAN packet loss, this multiplies 99th-percentile (P99) connection establishment latency by 4.5x (from 18.4 ms to 82.6 ms). • Side-Channel Immunity & Formal Verification: Zero secret-dependent execution paths were detected across 100 million DudeCT statistical traces (|t| < 1.28 << 4.5), confirming robust timing attack resistance. • Mitigation via RFC 8879: Implementing RFC 8879 certificate compression (Brotli/Zstandard) achieves a 58.4% reduction in certificate chain size, preventing TCP window exhaustion and restoring single-RTT connection establishment. Finally, we present an open-source, production-ready reference configuration for OpenSSL 3.4 and BoringSSL to facilitate secure and performant global migration to post-quantum internet infrastructure.
No takes yet. Share an insight, caveat, or question.
Kartik Kothalkar (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: