Methods paper describes a protocol for adversarial multi-AI review, demonstrating its application and implications.
This is a methods paper describing a protocol for adversarial multi-AI review — using multiple independent AI systems as falsification instruments against one's own thesis, with modular kill criteria, source/scope separation, explicit null reporting, and mandatory human adjudication. The protocol is demonstrated through a worked case involving a companion paper, "The Limits of Sacred Authority" (Zenodo, v1.5.1). The paper documents its own executed case honestly, including a doctrinal-search null result and the protocol's most consequential error — an over-read of that null as refutation, corrected by human adjudication against unanimous system agreement. It defines protocol-level falsification explicitly and separates it from refutation, using a seven-category outcome taxonomy. Section E documents four system failures observed during the executed case: semantic drift with unevidenced self-certification, assignment scope exceeded, wrong assignment answered, and fabricated/misdescribed legal citations. Failures are reported generically — no AI system or vendor is named in this paper — because the object is a transferable failure taxonomy, not an assessment of named products. The reproducibility boundary is stated explicitly: this is a narrative case report and reusable protocol, not a complete audit package. This upload includes three companion files: a public, redacted evidence-recovery packet supporting the evidentiary standing of the Section E observations (system identities withheld); the case-specific Sacred Authority Research Protocol, separated out because it is not reusable independently of subject matter; and a deposit manifest with SHA-256 hashes and version history. AI systems were used for structured review, adversarial critique, and drafting assistance throughout. They are not authors. The human author reviewed all sources, interpretations, and conclusions and takes sole responsibility for them.
No takes yet. Share an insight, caveat, or question.
QianJun Yu (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: