PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 17, 20250 citationsOpen Access

Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation

View Full Paper
XCXinrui ChenHZHongxing ZhangFZFanyi Zeng

Key Points

  • Prune&Comp addresses performance degradation by compensating for magnitude gaps due to layer removal.
  • With an iterative pruning strategy, Prune&Comp maintains 93.19% of the original model's performance while reducing perplexity nearly by half.
  • This technique offers zero runtime overhead while enhancing existing layer pruning metrics for large language models.
  • Applying the block influence metric significantly outperforms baseline models, demonstrating Prune&Comp's effectiveness.

Abstract

Layer pruning has emerged as a promising technique for compressing large language models (LLMs) while achieving acceleration proportional to the pruning ratio. In this work, we identify that removing any layer induces a significant magnitude gap in hidden states, resulting in substantial performance degradation. To address this issue, we propose Prune&Comp, a novel plug-and-play layer pruning scheme that leverages magnitude compensation to mitigate such gaps in a training-free manner. Specifically, we first estimate the magnitude gap caused by layer removal and then eliminate this gap by rescaling the remaining weights offline, with zero runtime overhead incurred. We further demonstrate the advantages of Prune&Comp through an iterative pruning strategy. When integrated with an iterative prune-and-compensate loop, Prune&Comp consistently enhances existing layer pruning metrics. For instance, when 5 layers of LLaMA-3-8B are pruned using the prevalent block influence metric, Prune&Comp nearly halves the perplexity and retains 93.19\% of the original model's question-answering performance, outperforming the baseline by 4.01%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2025) studied this question.

synapsesocial.com/papers/68f19f20de32064e504ddf59https://doi.org/10.48550/arxiv.2507.18212
Ask AI
Helpful
Bookmark
Share
View Full Paper