Experimentally verified SA-DSM-MADDPG enhances multi-agent control and training stability in pursuit-evasion scenarios, improving success rates by 22% compared to MADDPG.
Multi-UAV cooperative encirclement in pursuit–evasion scenarios requires effective coordination under dynamic inter-agent interactions, sparse task feedback, and obstacle-constrained motion. While MADDPG offers a practical CTDE framework for multi-agent continuous control, its direct application to cooperative encirclement still faces challenges in modeling time-varying teammate dependencies, selecting informative replay samples, and maintaining stable learning under delayed rewards. To address these challenges, we propose SA-DSM-MADDPG, an enhanced multi-agent deep deterministic policy gradient method that integrates the following: (i) a self-attention critic to model dynamic inter-agent relevance, (ii) a double-screened experience replay strategy combining prioritized sampling and relevance screening to improve replay quality, and (iii) curriculum learning with staged reward shaping to provide denser and more stable training signals. We evaluate the proposed method in 3v1 cooperative encirclement environments with static obstacles and varying initial conditions. Experimental results show that SA-DSM-MADDPG improves the success rate by approximately 22 percentage points over MADDPG and 35 percentage points over MAPPO, while also exhibiting faster convergence and better training stability.
No takes yet. Share an insight, caveat, or question.
Liang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: