PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 23, 20250 citationsOpen Access

CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning

View Full Paper
JFJinyuan FengCWChaopeng WeiTQTenghai Qiu

Key Points

  • CoMoE enhances the capacity of mixture-of-experts, promoting better specialization among experts in model training.
  • Experiments demonstrated improvements across various benchmarks, suggesting enhanced performance in heterogeneous datasets.
  • The method uses a contrastive objective that recovers information gaps between activated and inactivated experts.
  • The study emphasizes the importance of effective module training for optimal use of expert capacities.

Abstract

In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off between model capacity and computation overhead. However, current MoE variants fall short on heterogeneous datasets, ignoring the fact that experts may learn similar knowledge, resulting in the underutilization of MoE's capacity. In this paper, we propose Contrastive Representation for MoE (CoMoE), a novel method to promote modularization and specialization in MoE, where the experts are trained along with a contrastive objective by sampling from activated and inactivated experts in top-k routing. We demonstrate that such a contrastive objective recovers the mutual-information gap between inputs and the two types of experts. Experiments on several benchmarks and in multi-task settings demonstrate that CoMoE can consistently enhance MoE's capacity and promote modularization among the experts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Feng et al. (2025) studied this question.

synapsesocial.com/papers/68d4764731b076d99fa6e02fhttps://doi.org/10.48550/arxiv.2505.17553
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1HMoE: Heterogeneous Mixture of Experts for Language Modeling2024 · 2 citations
  2. 2MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts2024 · 14 citations
  3. 3Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast2024 · 1 citations
  4. 4Efficiently Editing Mixture-of-Experts Models with Compressed Experts2025
  5. 5Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning2025