Preventing lateral movement remains a central cybersecurity challenge even in environments designed according to Zero Trust principles. Although this paradigm reduces implicit trust and enforces continuous verification, its effectiveness ultimately depends on access-policy quality, identity-to-resource segmentation, and the ability to detect abusive chains built from seemingly legitimate permissions. In parallel, recent automated penetration-testing research has advanced through reinforcement learning, graph-based modeling, and simulation frameworks for exploring complex attack surfaces 1-4. Building on this state of the art, this article proposes a conceptual white-box penetration-testing framework for Zero-Trust architectures in which evolutionary algorithms perform global search over the internal blueprint of the environment, while reinforcement learning adaptively refines promising action sequences. The model assumes authorized defensive access to ZTNA policies, identity and privilege graphs, workload dependencies, and continuous-verification logs. Its fitness function is multi-objective and jointly considers success probability, stealth, and evasion rate. We argue that this combination may improve the identification of plausible lateral-movement routes and generate more useful remediation outputs, provided that it is applied in controlled environments with telemetry sufficiently faithful to the real system.
Marcelo Araujo (Sat,) studied this question.