PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20250 citationsOpen Access

AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning

View Full Paper
ZPZhenyu PanYZYiting ZhangZLZhuo Liu

Key Points

  • AdvEvo-MARL internalizes safety, reducing the attack-success rate significantly below traditional methods.
  • The method shows that optimizing both attackers and defenders enhances system resilience against threats.
  • Joint optimization preserves task accuracy while mitigating risks, demonstrating the framework's effectiveness.
  • In adversarial scenarios, AdvEvo-MARL consistently outperforms baseline safety protocols with lower costs.

Abstract

LLM-based multi-agent systems excel at planning, tool use, and role coordination, but their openness and interaction complexity also expose them to jailbreak, prompt-injection, and adversarial collaboration. Existing defenses fall into two lines: (i) self-verification that asks each agent to pre-filter unsafe instructions before execution, and (ii) external guard modules that police behaviors. The former often underperforms because a standalone agent lacks sufficient capacity to detect cross-agent unsafe chains and delegation-induced risks; the latter increases system overhead and creates a single-point-of-failure-once compromised, system-wide safety collapses, and adding more guards worsens cost and complexity. To solve these challenges, we propose AdvEvo-MARL, a co-evolutionary multi-agent reinforcement learning framework that internalizes safety into task agents. Rather than relying on external guards, AdvEvo-MARL jointly optimizes attackers (which synthesize evolving jailbreak prompts) and defenders (task agents trained to both accomplish their duties and resist attacks) in adversarial learning environments. To stabilize learning and foster cooperation, we introduce a public baseline for advantage estimation: agents within the same functional group share a group-level mean-return baseline, enabling lower-variance updates and stronger intra-group coordination. Across representative attack scenarios, AdvEvo-MARL consistently keeps attack-success rate (ASR) below 20%, whereas baselines reach up to 38.33%, while preserving-and sometimes improving-task accuracy (up to +3.67% on reasoning tasks). These results show that safety and utility can be jointly improved without relying on extra guard agents or added system overhead.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pan et al. (2025) studied this question.

synapsesocial.com/papers/68e25385d6d66a53c2474c1chttps://doi.org/10.48550/arxiv.2510.01586
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Co-Evolving Complexity: An Adversarial Framework for Automatic MARL Curricula2025
  2. 2Navigating Trust and Safety in Multi-Agent Reinforcement Learning: Attacks, Defenses, and Future Directions2024 · 1 citations
  3. 3Research on the evolution mechanism of multi-agent reinforcement learning marl simulation volleyball cooperative tactics2026
  4. 4Collaborative multi-agent reinforcement learning (C-MARL) for automated red teaming in large-scale heterogeneous networks2026
  5. 5Securing AI-Agentic Interactions via Multi-Agent Reinforcement Learning (MARL) with Secure Communication Protocols2024 · 1 citations