PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 20250 citationsOpen Access

Mixture of Experts Made Intrinsically Interpretable

View Full Paper
XYXingyi YangCVConstantin VenhoffAKAshkan Khakzar

Key Points

  • MoE-X achieves better interpretability while matching performance of dense language models.
  • Evaluation on chess and natural language tasks shows perplexity better than GPT-2.
  • Sparse activation within each expert enhances feature routing and interpretability objectives.
  • The model's architecture allows for efficient scaling while maintaining performance outcomes.

Abstract

Neurons in large language models often exhibit polysemanticity, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on post-hoc methods, we present MoE-X, a Mixture-of-Experts (MoE) language model designed to be intrinsically interpretable. Our approach is motivated by the observation that, in language models, wider networks with sparse activations are more likely to capture interpretable factors. However, directly training such large sparse networks is computationally prohibitive. MoE architectures offer a scalable alternative by activating only a subset of experts for any given input, inherently aligning with interpretability objectives. In MoE-X, we establish this connection by rewriting the MoE layer as an equivalent sparse, large MLP. This approach enables efficient scaling of the hidden size while maintaining sparsity. To further enhance interpretability, we enforce sparse activation within each expert and redesign the routing mechanism to prioritize experts with the highest activation sparsity. These designs ensure that only the most salient features are routed and processed by the experts. We evaluate MoE-X on chess and natural language tasks, showing that it achieves performance comparable to dense models while significantly improving interpretability. MoE-X achieves a perplexity better than GPT-2, with interpretability surpassing even sparse autoencoder (SAE) -based approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2025) studied this question.

synapsesocial.com/papers/68d90a0f41e1c178a14f6956https://doi.org/10.48550/arxiv.2503.07639
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Closer Look into Mixture-of-Experts in Large Language Models2024 · 4 citations
  2. 2ST-MoE: Designing Stable and Transferable Sparse Expert Models2022 · 50 citations
  3. 3MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models2025
  4. 4HMoE: Heterogeneous Mixture of Experts for Language Modeling2024 · 2 citations
  5. 5Evolving Sparse Spiking Mixture-of-Experts: A Unified Neuromorphic Language Modeling Framework2026