PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 24, 20240 citationsOpen Access

Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications

View Full Paper
YLYang LiCZChangsheng ZhaoHLHyungtak Lee

Key Points

Key points are not available for this paper at this time.

Abstract

Large language models (LLMs) significantly enhance the performance of various applications, but they are computationally intensive and energy-demanding. This makes it challenging to deploy them on devices with limited resources, such as personal computers and mobile/wearable devices, and results in substantial inference costs in resource-rich environments like cloud servers. To extend the use of LLMs, we introduce a low-rank decomposition approach to effectively compress these models, tailored to the requirements of specific applications. We observe that LLMs pretrained on general datasets contain many redundant components not needed for particular applications. Our method focuses on identifying and removing these redundant parts, retaining only the necessary elements for the target applications. Specifically, we represent the weight matrices of LLMs as a linear combination of base components. We then prune the irrelevant bases and enhance the model with new bases beneficial for specific applications. Deep compression results on the Llama 2-7b and -13B models, conducted on target applications including mathematical reasoning and code generation, show that our method significantly reduces model size while maintaining comparable accuracy to state-of-the-art low-rank compression techniques.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2024) studied this question.

synapsesocial.com/papers/68e68aacb6db6435876123dbhttps://doi.org/10.48550/arxiv.2405.15877
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization2024 · 2 citations
  2. 2Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models2024
  3. 3Data-free Weight Compress and Denoise for Large Language Models2024
  4. 4Designing Large Foundation Models for Efficient Training and Inference: A Survey2024 · 5 citations
  5. 5Efficient Compression of Large Language Models: A Case Study on Llama 2 with 13B Parameters2024 · 18 citations