PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20250 citationsOpen Access

DiEP: Adaptive Differentiable Expert Pruning for Mixture-of-Experts Compression

DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning

View Full Paper
Ask AI
Bookmark
Share

Authors

SBSikai BaiHLH. J. LiJZJie Zhang

Discussion

Loading...

Member takes

Overview

Adaptive pruning improves mixture-of-experts model efficiency in NLP tasks, retaining high performance.

Key Points

  • DiEP retains approximately 92% of performance with half the experts on Mixtral 8×7B and improves efficiency.
  • The proposed method outperforms existing pruning techniques by up to 7.1% on the MMLU dataset, enhancing model utility.
  • Non-uniform expert pruning addresses varying redundancy across layers, ensuring optimal outcome in model performance.
  • This approach transforms discrete search space into a continuous one, enabling effective gradient-based optimization.

Cite This Study

Bai et al. (2025) studied this question.

synapsesocial.com/papers/68e040eda99c246f578b3452https://doi.org/10.48550/arxiv.2509.16105
View Full Paper
Ask AI
Bookmark
Share