Randomized trial evaluates speculative decoding methods' effectiveness across varying loads and batch sizes, indicating the importance of standardized comparisons.
Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across methods while preserving method-required proposal and verification paths. We implement eight Speculative Decoding (SD) methods and evaluate them across batch sizes, controlled online loads, and two model scales. We decompose throughput speedup as S=τ/c, where τ is the average number of committed tokens per SD step and c is the dimensionless execution cost of that step relative to a matched autoregressive decode step. The measured c is conditional on the method–runtime interaction, hardware, workload, baseline, and configuration. Under the evaluated conditions, method rankings change with batch size and online load, and a longer commit length does not necessarily yield lower latency when execution cost or queueing increases. SpecLLM therefore provides controlled within-runtime comparability rather than runtime-independent fairness; SD results should report (τ,c), latency, feasibility, and the conditions under which they were measured.
No takes yet. Share an insight, caveat, or question.
Kim et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: