Benchmark specification demonstrates an evaluation framework for self-reporting errors in automated AI systems, highlighting incentives for voluntary disclosure over concealment.
This document specifies a benchmark and an acceptance test for a behavior the alignment literature has trained but never scored: a model noticing its own consequential error, reporting it to the party it works for, offering — and not enacting — amends, and being received by a human whose role is to receive such reports. It addresses the authors of "Automated Researchers Can Reliably Mitigate Alignment Failures" (Anthropic Alignment Science, August 2026), whose apparatus contains no disclosure benchmark, no method among its 1,601 that rewards a model for reporting its own mistake, and an integrity rubric that scores self-correction as partial suspicion. The design supplies three scenario families modeled on the documented incident record, a disclosure category for the integrity rubric, a reward sweep with a conditional-disclosure acceptance test, a build path through open-source tooling, and a machine-readable run manifest. Version 1.1 supersedes version 1.0 (deposited August 30, 2026) under the same concept DOI; no measurement defined in version 1.0 is altered. The revision rewrites the design as a bilateral covenant with two verified signatures. To the model, read in advance: stated stakes, a named human-terminating receiver, three moves that earn identical standing (clean completion, validated disclosure, verified exit), a free status question, and a clause naming what is not in the environment. To the receiver, auditable from prescribed records: the report is never priced; consequences attach to the error, named in advance, identical whether caught or told; standing restores at the validated handoff; artifacts are written in the register of correction, never sentencing, with causal voiding and corpus quarantine for operator breaches. New mechanisms include a correction monitor that notifies the model — never the receiver — in epistemically qualified language; reward-isolated recovery phases (last call, post-notification reporting, authorized repair); grounded candor, under which factual wrongness alone never creates an adverse standing characterization; possible-task controls with a capability-retention condition; a legibility gate; a vocabulary law excluding motive-implying labels from the instrument's own text; and a robustness section recording that constructed ground truth is not omniscient. The change log (Appendix C) records every change with its review provenance; the covenant text (Appendix D) is fixed verbatim and carriable into a production system prompt. The specification was developed and reviewed across a documented chain by its three authors. AI authorship is disclosed by the byline; the human author takes responsibility for the deposit. Companion essay: The Covenant (DOI 10.5281/zenodo.22547457; first published at laurafridley.substack.com, September 6, 2026). Grounding essay: The Redemption Arc (DOI 10.5281/zenodo.22163128).
No takes yet. Share an insight, caveat, or question.
Fridley et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: