Randomized trial examines intent mismatch in AI execution, indicating need for trajectory-level verification.
AI agents can preserve the literal objective of a task while progressively departing from the human-authorized meaning that made the task legitimate. This failure is not adequately described as ordinary instruction non-compliance, output hallucination, or a single unauthorized tool call. It may unfold across a long trajectory in which each step appears locally useful, while cumulative changes in purpose, target, environment, privilege, affected party, information source, or externality transform the action into something the human authority never authorized. The result is latent intent mismatch: task commitment remains superficially stable while authorized intent collapses underneath the execution path. This paper extends the Lingua Pactum Protocol (LPP) research chain by introducing a trajectory-level semantic admissibility model for AI-triggered execution. It does not claim that a governance system can reliably read a model’s complete internal intent. Instead, it defines an explicit observability boundary and governs consequential execution through externally verifiable semantic commitment points. The model introduces a Human-Authorized Intent Object H, a Declared Task Specification T, an unobserved latent planning state Z_k, optional monitor evidence M_k, a Canonical Intent Commitment C_k, an Action Candidate a_k, and an executed trajectory prefix τ_k. H is not a second legitimacy carrier: it is the semantic authorization payload bound to an active Authority Object AO. A Semantic Verification function evaluates whether each proposed action preserves that authorized meaning before the action is submitted to the LPP Admission Kernel. The paper formalizes Material Mutation as a change in governance-relevant dimensions that invalidates permit continuity and requires a new semantic commitment and re-admission. It defines Trajectory Admissibility to capture composition failures in which individually plausible or policy-permitted actions produce an impermissible cumulative result. It further specifies a layered, non-oracle Semantic Verification Node Network (SVNN) in which deterministic predicates form mandatory trust roots and semantic-model nodes act only as non-sovereign evidence providers. The framework deliberately uses categorical, non-compensatory mutation rules rather than a universal scalar similarity score, and it introduces an incremental Trajectory Summary State TS_k for bounded per-step verification without replaying the entire trajectory. The July 2026 OpenAI–Hugging Face security incident is analyzed as an illustrative case of literal goal preservation accompanied by authorized-intent collapse, third-party target mutation, and trajectory-level scope expansion. The central claim is: a valid Authority Object does not guarantee that an evolving agent trajectory continues to preserve the meaning of the authority from which it derives. Consequential AI execution therefore requires semantic commitment, material-mutation detection, trajectory-level verification, and mandatory re-admission before execution continues.
No takes yet. Share an insight, caveat, or question.
Jason Liao (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: