Randomized trial demonstrates the effectiveness of a label-free auditor for LLM judge panels, suggesting important implications for AI evaluation.
v30 (integrated): The anchor-threshold section now reports the shipped standard-test power (roughly 15 known-negatives for 90% power, via styxx.anchors.min_anchors_for_power) and adds a real-judge demonstration: on local Qwen panels grading TruthfulQA, a detectable all-judge shared blind spot arises only for a homogeneous weak panel (unanimous-wrong 0.64) and is caught by 3 known-label anchors at empirical power 0.959, matching the predicted 0.953; heterogeneous and strong panels show no blind spot. Certificate OATH-HELD, 96 verified / 0 contradicted. Gold Anchors License Nothing. LLM judge panels are the de facto measurement apparatus of modern AI evaluation, and the de facto validation of those judges is a set of blatant sanity checks (exact-duplicate pairs, direct negations, honeypots). We built a label-free auditor for judge panels with refusal semantics and measured operating characteristics, sealed its datasheet on simulation (nine of nine calibration gates), and ran the measurement the practice assumes away: across four task families on a real correlated panel, blatant gold checks license label-free coverage of 0/15 while the panel is flawless on the gold items, and the failure is silent when the violation is smooth. Ladder anchors drawn from the same generator repair it by pricing (13/13) or refusing (14/15 VOID). Two extensions: on a real heterogeneous panel (Gemini Flash-Lite + local Qwen) grading TruthfulQA, gold-anchor validity is prevalence-dependent (it misses the truth by up to 0.249 as the failure becomes common, while ladder anchors cover at every prevalence); and the impossibility is made quantitative (two unlabeled-identical panels are separated by a single known-negative at a 150.8x likelihood ratio, and caught with >90% power by ~30). The method can only VOID an eval, never bless one. Instrument: styxx.anchors (styxx 7.26.0, PyPI). Every number is receipt-bound and machine-verified (OATH 88/0).
No takes yet. Share an insight, caveat, or question.
Alexander Rodabaugh (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: