PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

LLaDA-MoE: A Sparse MoE Diffusion Language Model

View Full Paper
FZFengqi ZhuZYZebin YouYXYi Xing

Key Points

  • LLaDA-MoE demonstrates competitive performance among diffusion models, utilizing fewer active parameters during inference.
  • With a capacity of 7B parameters, LLaDA-MoE only activates 1.4B parameters, effectively reducing computational overhead.
  • The model surpasses previous benchmarks, proving the efficacy of a sparse MoE architecture in language model training.
  • Empirical evaluation highlights LLaDA-MoE's capabilities in knowledge understanding and code generation tasks.

Abstract

We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves competitive performance with significantly reduced computational overhead by maintaining a 7B-parameter capacity while activating only 1.4B parameters during inference. Our empirical evaluation reveals that LLaDA-MoE achieves state-of-the-art performance among diffusion language models with larger parameters, surpassing previous diffusion language models LLaDA, LLaDA 1.5, and Dream across multiple benchmarks. The instruct-tuned model LLaDA-MoE-7B-A1B-Instruct demonstrates capabilities comparable to Qwen2.5-3B-Instruct in knowledge understanding, code generation, mathematical reasoning, agent and alignment tasks, despite using fewer active parameters. Our results show that integrating a sparse MoE architecture into the training objective of masked diffusion language models still brings out MoE's strengths under efficient inference with few active parameters, and opens ample room for further exploration of diffusion language models. LLaDA-MoE models are available at Huggingface.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcce8d54a28a75cf1c23https://doi.org/10.48550/arxiv.2509.24389
Ask AI
Helpful
Bookmark
Share
View Full Paper