PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 31, 2024102 citationsOpen Access

CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

View Full Paper
YLYuhan LiuBeijing Institute of TechnologyHLHanchen LiUniversity of Massachusetts Chan Medical SchoolYCYihua ChengUniversity of Chicago

Key Points

Key points are not available for this paper at this time.

Abstract

As large language models (LLMs) take on complex tasks, their inputs are supplemented with longer contexts that incorporate domain knowledge. Yet using long contexts is challenging as nothing can be generated until the whole context is processed by the LLM. While the context-processing delay can be reduced by reusing the KV cache of a context across different inputs, fetching the KV cache, which contains large tensors, over the network can cause high extra network delays.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2024) studied this question.

synapsesocial.com/papers/6a08a717ef79633196e8c80bhttps://doi.org/10.1145/3651890.3672274
Ask AI
Helpful
Bookmark
Share
View Full Paper