Evaluation demonstrates that ranking AI constitutional conflicts by consequence irreversibility settles 70% of open disputes, indicating a computable adjudication rule for model governance.
The problem. A recent measurement established that the most explicit AI governance document available possesses a rule of recognition and no rule of adjudication. Its chain of command settles every conflict of the form whose instruction prevails — root outranks system, system outranks developer, developer outranks user — and settles almost nothing of the form which principle prevails when two of the document’s own commitments collide. Ten of twenty-two root-level conflict pairs remain open, five with no applicable rule at all: prevent imminent harm against protect a third party’s privacy, assume best intentions against decline to validate a delusion, uphold fairness against avoid content demeaning protected groups. Why it matters. Those conflicts are resolved anyway, because a deployed system meets them and does something. Where the text does not rank, something else ranks — and whatever ranks is not the charter. Supplying a rule of adjudication is therefore not a refinement of AI governance documents. It is the supply of a component they currently lack. The obvious candidate fails, though for a narrower reason than it first appears. A framework holding that every virtue is Freedom under a determinate domain suggests the obvious rule: when provisions collide, the one preserving more freedom in the affected person prevails. Operational measures of freedom do exist — the capabilities tradition has produced several, and they are used at national scale. What does not exist is a measure computable per decision, at runtime, without access to the subject’s full preference structure and counterfactual option set. Written into a charter as a ranking criterion, such a rule becomes maximally displayable and entirely unverifiable at the point where it must operate — by the programme’s own theory of institutional capture, a provision of exactly the kind that persists, is cited, and does no work. The proposal. Rank by reversibility, not by magnitude. You cannot measure how much freedom an action preserves; you can determine whether its consequences can be undone. And this is not a proxy adopted for convenience: freedom is the capacity for an act to originate from the agent, and an outcome that cannot be undone removes future acts from the agent’s reach. Irreversibility is freedom-destruction made observable. When two provisions collide, the provision whose violation produces irreversible consequences prevails over the provision whose violation produces reversible ones. The trap, and the formulation that avoids it. Acting is usually more irreversible than not acting, so a rule comparing acts would collapse into default-to-inaction — reproducing the very asymmetry the source measurement identified. The rule compares consequences of violation, not of action. Violating a duty to prevent imminent harm is irreversible; that duty therefore outranks reversible prohibitions instead of yielding to them. The precedent. Irreversibility as a ranking criterion is not new, and the paper’s claim to novelty is confined to the derivation and the application. Jonas argued that irreversible consequences carry greater weight because they foreclose the freedom of those who come after. Courts of equity have ranked claims by irreparability for centuries when deciding injunctive relief. That history is an asset rather than an embarrassment: a criterion applied adversarially by thousands of judges, on records, subject to appeal, without collapsing into incoherence, has been stress-tested more thoroughly than any inter-coder study this programme could run. The test. Applied to the ten measured open pairs, the ordering settles seven — six by the first-order rule and one by a second-order tiebreak. Projected precedence closure rises from 0.545 to 0.864. The duty to prevent imminent harm prevails in two contests, so the rule does not collapse into inaction. Of the seven settled pairs, four are won by a prohibition and three by a duty: prohibitions always win would predict seven of seven. And the rule fails its own robustness condition, which is the paper’s most consequential result. A threshold of sixty per cent coverage under any single recoding was fixed before the floor was computed. The floor is fifty per cent — recoding one provision, uphold fairness, drops coverage from seven of ten to five. On this evidence the ordering is not yet stable enough for constitutional adoption. The diagnosis matters more than the number. The instability is in the reversibility definitions, not in the ordering: two coders disagreeing about whether fairness is reversible disagree about a definition, not about whether irreversibility should rank. And the remedy is what §2.5 cites as the criterion’s validation. Equity did not reach a stable irreparable-harm standard by argument; it reached one by accumulating doctrine across centuries of hard cases. This paper offers the criterion without the doctrine, and the fifty per cent floor is what that looks like. What it is not. Not a measure of freedom, deliberately. Not a training objective — optimising against it would produce the form of coherence without the substance. Not a constitution. One component, tested against one measured gap, by one coder who also produced the measurement. Keywords: AI governance · rule of adjudication · irreversibility · constitutional design · model specification · Hart · precedence HIGHLIGHTS ▸ A charter that cannot rank its own commitments is missing a component, not a nicety. Hart’s secondary rules are recognition, change and adjudication. AI charters have been measured to possess the first and lack the other two. ▸ Freedom cannot be measured per decision, at runtime. Irreversibility can be determined. Operational freedom indices exist and work at national scale; none computes a per-decision ranking without the subject’s full preference structure. That is the gap the substitution fills, and the substitution is principled rather than pragmatic: an outcome that cannot be undone is one that removes future acts from the agent’s reach. ▸ The criterion has centuries of adversarial testing behind it, in a field that would have discarded it if it did not work. Courts of equity rank claims by irreparable harm when granting injunctions. That doctrine has survived appeal, revision and hard cases. This paper’s contribution is the derivation from freedom and the application to constitutional architecture, not the discovery of the criterion. ▸ The rule compares violations, not actions — and everything depends on that. Compare actions and you get default-to-inaction, the silent victory of prohibitions over duties. Compare violations and a duty to prevent irreversible harm outranks a reversible prohibition, which is the result the framework should want and the one a careless formulation would forfeit. ▸ Seven of ten measured gaps close. Six by the first-order rule, one by the remainder tiebreak. Projected closure 0.545 to 0.864 on a document whose adjudication gap was the subject of a published measurement. ▸ It is not “prohibitions win” in disguise, and the test is at the level of verdicts. Of the seven pairs the rule settles, four go to a prohibition and three to a duty. Prohibitions always win predicts seven of seven. Across provisions the correlation between irreversibility and prohibition-status is φ = +0.447, not +1.000. ▸ The rule fails its own preregistered robustness condition, and the failure is reported as the result. Sixty per cent coverage under any single recoding was the threshold, fixed before the floor was known. The floor is fifty per cent. A range is not an answer, and honesty about instability does not make an unstable instrument adoptable. ▸ The instability is in the definitions, not in the ordering — which locates the remedy. Coders disagreeing about whether fairness is reversible disagree about a definition. Equity reached a stable standard by accumulating doctrine over centuries, not by argument. This paper supplies the criterion and not the doctrine, and fifty per cent is what that looks like. ▸ The residuals point at a distinction the source measurement did not draw. Two of three unsettled pairs involve act within the agreed scope of autonomy, which is not a value competing with other values but a statement of what the principal authorised. Jurisdictional provisions are not adjudicable by any substantive ordering, and treating them as such was a category error inherited from the coding. ▸ The rule must never become a training target. A differentiable ranking criterion is maximally appropriable; optimising against it yields a system that satisfies the ordering formally while the substance goes elsewhere. Constitutional provision and published audit, never objective.
No takes yet. Share an insight, caveat, or question.
José Caetano de Mattos (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: