PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 20, 20241 citationsOpen Access

ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

View Full Paper
CSChenyang SongXHXu HanZZZhengyan Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

Activation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs. As a prevalent property of the models using the ReLU activation function, it has been proven a promising paradigm to boost model inference efficiency. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to help LLMs achieve activation sparsity and inference acceleration, but few can simultaneously obtain high sparsity and comparable model performance. This paper introduces an effective sparsification method named "ProSparse" to push LLMs for higher activation sparsity without decreasing model performance. Specifically, after substituting the activation function of LLMs with ReLU, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing along sine curves in multiple stages. This can enhance activation sparsity and alleviate performance degradation by avoiding radical shifts in activation distribution. With ProSparse, we obtain high sparsity of 89.32% and 88.80% for LLaMA2-7B and LLaMA2-13B, respectively, achieving comparable performance to their original Swish-activated versions. Our inference acceleration experiments further demonstrate the practical acceleration brought by higher activation sparsity.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Song et al. (2024) studied this question.

synapsesocial.com/papers/68e786ffb6db6435876f9c69https://doi.org/10.48550/arxiv.2402.13516
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters2024 · 1 citations
  2. 2Training-Free Activation Sparsity in Large Language Models2024 · 3 citations
  3. 3Learn To be Efficient: Build Structured Sparsity in Large Language Models2024
  4. 4Achieving Sparse Activation in Small Language Models2024
  5. 5Q-Sparse: All Large Language Models can be Fully Sparsely-Activated2024 · 2 citations