PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 2, 2026Applied Energy3 citationsOpen Access

Mixture-of-experts based multi-critic deep reinforcement learning for sustainable management of data center microgrids

View Full Paper
QLQiong LiuZPZhanhua PanYGYe Guo

Key Points

  • The research aims to enhance the sustainable management of data center microgrids using advanced reinforcement learning techniques.
  • Formulated the management issue as a multi-reward Markov Decision Process (MDP)
  • Developed a multi-critic deep reinforcement learning framework with dedicated critics for each reward
  • Implemented a multi-head architecture based on a mixture-of-experts structure
  • The proposed method significantly reduces overall energy consumption in data center microgrids
  • Carbon emissions are also notably lowered compared to existing methods
  • Performance surpasses both state-of-the-art deep reinforcement learning and traditional rule-based approaches

Abstract

The rapid expansion of data centers (DCs) has led to substantial increases in global energy consumption and carbon emissions. Moreover, the strong coupling among workload scheduling, IT, cooling, and energy subsystems makes sustainable DC microgrids highly complex, especially given the difficulty of accurate modeling and the presence of significant uncertainties and rapid dynamics. To address these challenges, this paper proposes a multi-critic deep reinforcement learning (DRL) approach based on a mixture-of-experts (MoE) with multi-head architecture. The problem is formulated as a multi-reward Markov Decision Process (MDP) with independent reward signals for workload efficiency, energy consumption, and carbon footprint, providing richer feedback compared to traditional single-reward formulations. In the proposed approach, each reward is assigned a dedicated critic, thereby avoiding the conflicts and interference that arise when a single critic attempts to learn multiple competing rewards simultaneously. The shared MoE foundation with the multi-head architecture enables each critic head to focus on learning the value function for its specific optimization target, while the integrated adaptive gating mechanisms facilitate dynamic leveraging of both common and task-specific knowledge. This architecture improves learning stability, accelerates convergence, and enhances adaptability to complex, multi-objective learning tasks. Extensive experiments on a DC microgrid show that the proposed method outperforms state-of-the-art DRL and rule-based baselines in reducing the overall energy consumption and carbon emissions. All codes can be found at https://github.com/ikelq/Mixture-of-experts-based-Multi-Critic-DRL-for-Sustainable-Management-of-Data-Center-Microgrids . • Formulate the sustainable management of a DC microgrid as a multi-reward MDP. • Propose a multi-critic DRL framework with each reward assigned to a dedicated critic. • Design a multi-head architecture with a shared MoE structure within the multi-critic DRL framework.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69a52920f1e85e5c73bf07c7https://doi.org/10.1016/j.apenergy.2026.127561
Ask AI
Helpful
Bookmark
Share
View Full Paper