DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
Authors
Loading...
Adaptive pruning improves mixture-of-experts model efficiency in NLP tasks, retaining high performance.
Bai et al. (2025) studied this question.