Methodological framework demonstrates verifiable kill-switch containment for autonomous AI, highlighting pathways for empirical regulatory auditing.
Existing regulatory and standards guidance on human oversight of autonomous AI systems, including the EU AI Act’s human oversight requirements for high-risk systems [1], NIST’s AI Risk Management Framework [2], ISO/IEC 42001 [3], and Singapore’s proposed MAS Guidelines on AI Risk Management [4] and IMDA Model AI Governance Framework for Agentic AI [5], establishes that oversight and intervention capability must exist. None of these frameworks specify how an independent auditor verifies that a stop mechanism actually works under the conditions where it is needed. This paper proposes a conformance-based audit approach: a defined Target of Evaluation, five testable control families (trigger recognition, authority, cessation, latency, and failure resilience), and evidence requirements attached to each that distinguish measured system behavior from procedural self-attestation. The paper also specifies an example auditable criterion and a machine-readable control representation. The intent is to convert descriptive containment guidance into auditable criteria consistent with established independent third-party audit methodologies for AI systems, and to propose a companion vendor-neutral assurance harness, connecting to deployments through pluggable adapters, capable of generating the evidence such criteria would require.
No takes yet. Share an insight, caveat, or question.
Naveen Sundaresan (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: