Randomized trial measures the effectiveness of a label-free auditor in AI evaluation, highlighting implications for quality assurance.
Gold Anchors License Nothing. LLM judge panels are the de facto measurement apparatus of modern AI evaluation, and the de facto validation of those judges is a set of blatant sanity checks (exact-duplicate pairs, direct negations, honeypots). We built a label-free auditor for judge panels with refusal semantics and measured operating characteristics, sealed its datasheet on simulation (nine of nine calibration gates), and ran the measurement the practice assumes away: across four task families on a real correlated panel, blatant gold checks license label-free coverage of 0/15 while the panel is flawless on the gold items, and the failure is silent when the violation is smooth. Ladder anchors drawn from the same generator repair it by pricing (13/13) or refusing (14/15 VOID). Two extensions: on a real heterogeneous panel (Gemini Flash-Lite + local Qwen) grading TruthfulQA, gold-anchor validity is prevalence-dependent (it misses the truth by up to 0.249 as the failure becomes common, while ladder anchors cover at every prevalence); and the impossibility is made quantitative (two unlabeled-identical panels are separated by a single known-negative at a 150.8x likelihood ratio, and caught with >90% power by ~30). The method can only VOID an eval, never bless one. Instrument: styxx.anchors (styxx 7.26.0, PyPI). Every number is receipt-bound and machine-verified (OATH 88/0).
No takes yet. Share an insight, caveat, or question.
Alexander Rodabaugh (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: