PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 11, 2025Electronics2 citationsOpen Access

MoE-World: A Mixture-of-Experts Architecture for Multi-Task World Models

View Full Paper
TCTang CongYLYuang LiuYWYanhan Wu

Key Points

  • The aim is to improve multi-task learning in world models by addressing the seesaw phenomenon during training.
  • Proposed Mixture-of-Experts architecture (MoE-World) integrates Transformer blocks with MoE layers.
  • Gating mechanisms are implemented using multilayer perceptrons.
  • Experiments conducted on standard benchmarks to evaluate the effectiveness.
  • MoE-World significantly mitigates the seesaw phenomenon.
  • Achieves competitive performance in world model's reward metrics.
  • Enhances both accuracy and efficiency of multi-task learning.

Abstract

World models are currently a mainstream approach in model-based deep reinforcement learning. Given the widespread use of Transformers in sequence modeling, they have provided substantial support for world models. However, world models often face the challenge of the seesaw phenomenon during training, as predicting transitions, rewards, and terminations is fundamentally a form of multi-task learning. To address this issue, we propose a Mixture-of-Experts-based world model (MoE-World), a novel architecture designed for multi-task learning in world models. The framework integrates Transformer blocks organized as mixture-of-experts (MoE) layers, with gating mechanisms implemented using multilayer perceptrons. Experiments on standard benchmarks demonstrate that it can significantly mitigate the seesaw phenomenon and achieve competitive performance on the world model’s reward metrics. Further analysis confirms that the proposed architecture enhances both the accuracy and efficiency of multi-task learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cong et al. (2025) studied this question.

synapsesocial.com/papers/69401b172d562116f28f74e9https://doi.org/10.3390/electronics14244884
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LIMT: Language-Informed Multi-Task Visual World Models2024
  2. 2Multi-world Model in Continual Reinforcement Learning2024
  3. 3Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models2024
  4. 4The Intersection of Modular Architectures and Scalable AI Systems2025
  5. 5PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning2024