PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 20241 citationsOpen Access

Accelerating Speculative Decoding using Dynamic Speculation Length

View Full Paper
JMJonathan MamouOPOren PeregDKDaniel Korat

Key Points

Key points are not available for this paper at this time.

Abstract

Speculative decoding is a promising method for reducing the inference latency of large language models. The effectiveness of the method depends on the speculation length (SL) - the number of tokens generated by the draft model at each iteration. The vast majority of speculative decoding approaches use the same SL for all iterations. In this work, we show that this practice is suboptimal. We introduce DISCO, a DynamIc SpeCulation length Optimization method that uses a classifier to dynamically adjust the SL at each iteration, while provably preserving the decoding quality. Experiments with four benchmarks demonstrate average speedup gains of 10.3% relative to our best baselines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mamou et al. (2024) studied this question.

synapsesocial.com/papers/68e6b3acb6db643587634c2chttps://doi.org/10.48550/arxiv.2405.04304
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput2024
  2. 2SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths2024 · 1 citations
  3. 3Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion2024
  4. 4SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences2026
  5. 5The Disparate Impacts of Speculative Decoding2025