Evaluation study reveals high internal repeatability but low inter-evaluator agreement in AI judging pipelines, highlighting the necessity of structured provenance profiles for reliable audits.
LLM-as-a-judge systems make evaluation scalable, but repeatability, agreement, human ratification, and downstream authority are different properties of an evaluation process. I examine those distinctions through two version-controlled evaluation programs and a set of naturalistic audit cases. In a 600-case program with four repeated evaluator runs, 475 of 480 primary cases were identical across all four runs, yet repeated-decision agreement with a synthetic reference was 1,488/1,920 (0.775). All 120 held-out cases were four-of-four stable, while 20 systematically disagreed with the synthetic reference. Two procedurally independent, non-cross-exposed model reviewers initially aligned 20/20 with the reference, but later recovery of the exact evaluator rubric showed that those reviewers had judged a broader construct than the stored metric measured. The historical review objects remained evidence for the construct they actually assessed; the supported inferential scope narrowed. A second program used 24 blinded comparison cases. Agreement between a primary machine judge and a model-originated review object that I subsequently reviewed and ratified was 8/24 (0.333; Wilson 95% CI 0.180-0.533). A burden score was also semantically inverted across evaluator surfaces, invalidating a raw numerical comparison until an explicit comparison transform was declared. These results do not establish human or model superiority. They motivate an evaluator-level provenance profile that records role, exposure, construct, comparator, peer exposure, output type, and downstream decision authority so that later audits can revise interpretation without rewriting historical evaluator outputs.
No takes yet. Share an insight, caveat, or question.
Logan Davis (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: