Effective policy optimization in multiagent reinforcement learning (MARL) necessitates extensive exploration of high-dimensional state-action spaces. However, such exploration may not only trigger unsafe states but also compromise system stability, posing significant challenges for deployment in safety-critical systems. To address this challenge, this article proposes a safety-stability layer that integrates robust control barrier functions (RCBFs) and input-to-state stable control Lyapunov functions (ISS-CLFs) for multiagent systems operating in unknown environments with uncertain dynamics. Furthermore, by integrating safety-stability constraints with a MARL framework, during the training phase, we exclusively focus on goal-reaching objectives to expand the policy network's exploration space, while in the deployment phase, policy outputs are filtered through a real-time safety-stability layer. In addition, an event-triggered mechanism for action compensation calculation is designed based on safety condition assessments to conserve computational resources. Finally, the effectiveness of the proposed method is validated through simulation experiments in dynamic multiunicycle environments. The results demonstrate that our approach not only ensures strict adherence to safety constraints but also significantly enhances the task execution efficiency of multiagent systems.
Jia et al. (Thu,) studied this question.