Multi-agent pursuit-evasion has significant applications in military, transportation, and industrial sectors. This task faces dual challenges of uncertain swarm scale and environmental uncertainty within unstructured dynamic environments. To address these, we propose a hierarchical reinforcement learning framework based on permutation invariance to balance pursuit efficiency and collision avoidance safety. First, in the evaluation stage, we introduce a hybrid feature aggregation mechanism based on the DeepSets structure and a predicted intercept point auxiliary task. By extracting permutation-invariant group features and introducing kinematic priors, this approach achieves zero-shot transfer and rapid convergence of the method across different swarm scales. Second, in the execution stage, we construct a residual gating architecture based on task decomposition. This architecture utilizes a frozen basic tracking stream to handle the global game and employs a conditioned residual stream with a dynamic gating mechanism to manage local obstacle avoidance, effectively resolving the multi-objective conflict problem under sparse rewards. Finally, extensive simulations and physical experiments demonstrate that the proposed method, after training on small-scale swarms, achieves robust zero-shot transfer to medium and large-scale swarms, completing efficient and stable pursuit tasks. Furthermore, the method was successfully transferred to a physical platform, completing tasks smoothly even in the presence of faulty agents, thereby validating its applicability in the real world.
Yang et al. (Wed,) studied this question.