Empirical benchmark demonstrates tail-latency trade-offs between GPU decoding schedules in 5G NR, indicating batch size strongly impacts worst-case delays.
Key Points
Characterize and compare tail latency across GPU-accelerated 5G NR LDPC decoding schedules, batch granularities, and timing boundaries.
Benchmarked FP32-layered and flooding decoders across 600,000 schedule observations on a single dynamically clocked WSL2 GPU.
Evaluated performance across three timing boundaries (full, decode-only, and transfer-only) under varying signal conditions (2-dB and 4-dB cells) and batch sizes (64 to 128).
Flooding decoders exhibited higher P50 and P99 latencies in 15 full-boundary matched-cap comparisons, whereas layered decoders reached first syndrome satisfaction earlier in every estimable 2- and 4-dB cell.
Full-boundary J99 increased by 234.7 μs when scaling from batch 64 to 128 (pointwise descriptive 95% percentile interval, 196.1–281.5 μs) over six sessions, while J99 and R99 showed no uniform schedule ordering.