PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 27, 20240 citationsOpen Access

Enhancing and Accelerating Large Language Models via Instruction-Aware Contextual Compression

View Full Paper
HHHaowen HouFMFei MaBBBinwen Bai

Key Points

Key points are not available for this paper at this time.

Abstract

Large Language Models (LLMs) have garnered widespread attention due to their remarkable performance across various tasks. However, to mitigate the issue of hallucinations, LLMs often incorporate retrieval-augmented pipeline to provide them with rich external knowledge and context. Nevertheless, challenges stem from inaccurate and coarse-grained context retrieved from the retriever. Supplying irrelevant context to the LLMs can result in poorer responses, increased inference latency, and higher costs. This paper introduces a method called Instruction-Aware Contextual Compression, which filters out less informative content, thereby accelerating and enhancing the use of LLMs. The experimental results demonstrate that Instruction-Aware Contextual Compression notably reduces memory consumption and minimizes generation latency while maintaining performance levels comparable to those achieved with the use of the full context. Specifically, we achieved a 50% reduction in context-related costs, resulting in a 5% reduction in inference memory usage and a 2.2-fold increase in inference speed, with only a minor drop of 0.047 in Rouge-1. These findings suggest that our method strikes an effective balance between efficiency and performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hou et al. (2024) studied this question.

synapsesocial.com/papers/68e5adbeb6db643587547177https://doi.org/10.48550/arxiv.2408.15491
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1In-Context Former: Lightning-fast Compressing Context for Large Language Model2024
  2. 2LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models2023 · 129 citations
  3. 3Adapting LLMs for Efficient Context Processing through Soft Prompt Compression2024 · 4 citations
  4. 4Contemporary Model Compression on Large Language Models Inference2024 · 5 citations
  5. 5Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention2025