Modern enterprise networks have grown into large-scale, heterogeneous environments spanning tens of thousands of nodes across diverse operating systems, including Linux servers, Windows Active Directory domains, and Internet-of-Things (IoT) clusters. Traditional red-teaming methodologies are manual, time-consuming, and inherently point-in-time, failing to keep pace with the dynamic nature of contemporary cyber-threats. Existing automated tools such as CALDERA operate on static rule-bases and collapse against moving-target defenses. Single-agent reinforcement learning (RL) approaches suffer acutely from the curse of dimensionality as network scale increases. To bridge this gap, we propose C-MARL a Collaborative Multi-Agent Reinforcement Learning framework in which five heterogeneous specialized agents jointly learn to navigate, exploit, and exfiltrate data from a 10,000-node simulated enterprise network. Our framework introduces three principal technical innovations: (i) a Universal Vectorized Adversarial Space (UVAS) that normalises observations across disparate OS types into a unified agent-agnostic state representation; (ii) a Gated Communication protocol that constrains agent-to-agent messaging to events where information gain exceeds a learnable threshold, with gradient flow through the binary gate handled via Gumbel-Softmax relaxation, reducing communication overhead by 47.5% compared with full-communication baselines; and (iii) a Hierarchical Credit Assignment mechanism based on exact Shapley values computed over all 2 5 = 32 coalitions, equitably distributing terminal rewards across the cooperative attack chain. RL training used an abstracted Gym-level state representation for scalability; GNS3/QEMU network emulation was reserved for final validation. C-MARL achieves an Attack Success Rate (ASR) of 89.2%, a Mean Time to Compromise (MTTC) of 38.4 steps, and a Vulnerability Coverage of 89.2% over 17,893 unique CVEs—outperforming single-agent RL, an Expert Rule-Based System (ERBS), and random-search baselines on all metrics. These results demonstrate that collaborative, specialized agents can substantially reduce red-teaming cycle time while maintaining stealth and operational relevance.
Alluhaidan et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: