PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 20240 citationsOpen Access

Accelerating Production LLMs with Combined Token/Embedding Speculators

View Full Paper
DWDavis WertheimerJRJoshua RosenkranzTPT. A. Parnell

Key Points

Key points are not available for this paper at this time.

Abstract

This technical report describes the design and training of novel speculative decoding draft models, for accelerating the inference speeds of large language models in a production environment. By conditioning draft predictions on both context vectors and sampled tokens, we can train our speculators to efficiently predict high-quality n-grams, which the base model then accepts or rejects. This allows us to effectively predict multiple tokens per inference forward pass, accelerating wall-clock inference speeds of highly optimized base model implementations by a factor of 2-3x. We explore these initial results and describe next steps for further improvements.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wertheimer et al. (2024) studied this question.

synapsesocial.com/papers/68e6d2d6b6db643587650468https://doi.org/10.48550/arxiv.2404.19124
Ask AI
Helpful
Bookmark
Share
View Full Paper