Autonomous AI agents are given objectives rather than procedures and search for their own paths to them. Reported 2026 incidents show that an agent's effective reach can exceed what anyone authorized, both where an intended limit was absent and where it was present and defeated. This paper specifies an architecture for keeping that reach within granted authority. The requirement is an invariant over sets of actions, E(s) ⊆ A(G(s), s), that must hold in every reachable state. Violations fall into three classes: unmediated reach, compromised decision, and authority expansion. Against a threat model in which the authorized subject is itself a search process, we derive six security properties. The reference architecture separates an untrusted agent domain from a control domain and issues taskbound, default-deny, attenuating grants. It enforces them at effect boundaries and adds independent resource-side verification and consequence ceilings for high-consequence resources. It also verifies environment reach and ties response bounds to time-to-effect. Independence between controls is claimed only against a named adversary and a dependency set. In an analytical replay of four reported incidents, two are prevented under stated assumptions, one is contained or detected, and one is not necessarily prevented. The architecture composes established protection mechanisms rather than replacing them. It is not formally verified, and the paper states where its claims stop.
No takes yet. Share an insight, caveat, or question.
Wilson Robert Wayne (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: