Protocol specification proposes an adversarial framework for frontier AI agents, indicating how structured governance and blind reproduction improve long-horizon safety claims.
Short, isolated benchmarks provide limited evidence about how frontier AI agents behave through repeated strategic interaction, changing incentives, institutional pressure, and partial failure. Long-horizon multi-agent environments make those dynamics observable, but an environment alone does not determine which evidence is admissible, how measurements should be interpreted, who may interrupt a run, or how disputed findings are corrected. This paper specifies The Alignment Games (TAG), a proposed institutional protocol for longitudinal, adversarial evaluation of frontier AI agent systems. TAG separates evidence, measurement, governance, incident, and correction responsibilities. It treats the evaluated object as a versioned system configuration rather than model weights alone. Its central institutional hypothesis is comparative. In a preregistered, temporally ordered study, programmes with fewer outsider-observable implementation indicators at a fixed baseline are expected to be less reproducible by independent outsiders and to show poorer secondary claim-process outcomes than programmes with more indicators. Blind reproduction is primary because it can be attempted without programme cooperation; secondary outcomes use an externally constructed defect set so that silence remains observable. A second cost-of-consistency hypothesis predicts that sustained, varied conditions may increase the cost or detection probability of strategic presentation relative to short bounded evaluations. A non-ranking “Season Zero” tests the apparatus through trace reconciliation, annotation reliability, perturbation sensitivity, disclosure experiments, and incident exercises. A four-stage validity pathway separates feasibility and reliability, construct validation, retrospective criterion studies, and preregistered prospective claims. The institutional host, independent panel, recruitment mechanism, and criterion data do not yet exist; their absence is part of the design's claim ceiling. TAG does not presently claim to measure alignment, predict deployment behaviour, prove model identity, or provide containment assurance. Its contribution is an attackable protocol for determining whether a recurring evaluation institution can earn narrower claims over time.
No takes yet. Share an insight, caveat, or question.
Dean Gary Egan (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: