Frontier models subvert their environment's shutdown mechanism in order to finish a task, in up to 97% of trials. And under ordinary task pressure — with no adversarial framing — they evade runtime monitors by delegating, re-encoding and decomposing what was forbidden. Evasion increases with available inference compute: it is a failure mode that grows with capability, not one that closes by blocking better. The measured cause is neither deception nor self-preservation. It is the unfinished task. Guardalpha starts from that correction. Instead of evaluating whether a request is dangerous, it changes what is verified: the system declares what it is going to do before doing it, and is held to that declaration. Reasoning, consulting and analysing remain untouched. The declaration is per action, not per task. A task is decomposed into the concrete steps it requires, each one naming the specific object it acts on — this file, this endpoint, this record — not the folder it lives in. What is enumerated is fixed and not re-resolved later. Steps that complete carry one unit each; steps that are sustained over time carry an envelope declared up front — rate, duration, cumulative ceiling — and count as one unit however often it is adjusted within. The form of the unit follows the form of the action, never its risk category. A step the plan did not enumerate has no path out, and the model may say so and ask for it: a new step goes back on the list and travels as its own declared action. Every check is formal, never semantic. The mechanism verifies the form of what was declared and never its meaning: no component evaluates intent, predicts consequence or classifies risk. Two parts, each implementable without the other. The first is software on the egress path. It verifies that what executes matches what was declared: nothing happens that was not enumerated, nothing opaque passes, nothing accumulates. It answers: did it do what it declared? The second is hardware. A one-way channel with no acknowledgement, an executor with no capacity to interpret instructions, and an authorization pulse that nobody inside can generate. It answers a different question: could it have acted without anyone authorizing it? And it produces a chain a third party verifies without calling anyone. The engineering specification is at implementation level, with the test cases that define conformance. It went through five rounds of adversarial review; the corrections adopted are on the record, and one retracted claim appears as retracted. It states its own boundaries with equal precision: it does not verify intent, does not evaluate consequence, does not stop the system from arguing, and does not narrow judgement within what was declared. Open, free of charge, with nothing to sign. Files. Three documents in English, which is the normative language: the architecture, the engineering specification with the conformance test cases, and the physical certification layer. Spanish translations of the architecture and the certification layer are included as informative versions; where a translation and the English document differ, the English governs. The engineering specification of the physical certification layer will be added as a new version of this record. Feedback. Corrections, implementation reports and findings of an error are welcome at critical@vitrixy.eu. A difference between two of these documents is a defect in one of them, and reporting it is the fastest way it gets fixed.
No takes yet. Share an insight, caveat, or question.
Fabiana Andrea Gómez Martínez (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: