PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 6, 20240 citationsOpen Access

Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment

View Full Paper
AAAbhinav AgarwallaAGAbhay GuptaAMAlexandre Carriconde Marques

Key Points

  • Sparse foundational models can achieve full accuracy recovery at up to 70% sparsity, improving performance significantly.
  • Training on the SlimPajama and The Stack datasets led to an acceleration on Cerebras CS-3 chips, matching theoretical models.
  • Inference speedups of up to 3x on CPUs and 1.7x on GPUs highlight the effectiveness of Neural Magic's DeepSparse engine and nm-vllm approach for fast computations with sparse models in diverse tasks including code generation and summarization. This suggests that further optimizations can be realized through quantization techniques.

Abstract

Large language models (LLMs) have revolutionized Natural Language Processing (NLP), but their size creates computational bottlenecks. We introduce a novel approach to create accurate, sparse foundational versions of performant LLMs that achieve full accuracy recovery for fine-tuning tasks at up to 70% sparsity. We achieve this for the LLaMA-2 7B model by combining the SparseGPT one-shot pruning method and sparse pretraining of those models on a subset of the SlimPajama dataset mixed with a Python subset of The Stack dataset. We exhibit training acceleration due to sparsity on Cerebras CS-3 chips that closely matches theoretical scaling. In addition, we establish inference acceleration of up to 3x on CPUs by utilizing Neural Magic's DeepSparse engine and 1.7x on GPUs through Neural Magic's nm-vllm engine. The above gains are realized via sparsity alone, thus enabling further gains through additional use of quantization. Specifically, we show a total speedup on CPUs for sparse-quantized LLaMA models of up to 8.6x. We demonstrate these results across diverse, challenging tasks, including chat, instruction following, code generation, arithmetic reasoning, and summarization to prove their generality. This work paves the way for rapidly creating smaller and faster LLMs without sacrificing accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Agarwalla et al. (2024) studied this question.

synapsesocial.com/papers/68e6b6eab6db64358763862fhttps://doi.org/10.48550/arxiv.2405.03594
Ask AI
Helpful
Bookmark
Share
View Full Paper