Controlled evaluation reveals negligible performance differences in automated decision-state grading, highlighting how undetected scorer defects can compromise benchmark results.
A controlled 288-run evaluation of one coded decision-state slice produced a reproducible as-scored disposition under its original scoring contract. The study contained 144 runs per arm and 864 completed turns, with blind grading and a project-recorded logistic mixed-effects analysis whose frozen code is now directly reproducible. Observed governed-task success was 19/144 in Arm A and 23/144 in Arm B. Marginal mixed-model estimates were 0.1316 and 0.1591, respectively, for a +2.75 percentage-point contrast (95% CI -4.12 to +9.62 pp). A recovered project custody record states a +10 percentage-point practical-effect criterion with a positive lower confidence bound; no independently inspectable pre-result artifact establishing that criterion was recovered. Under that project-recorded criterion, the as-scored coded slice did not satisfy the stated practical-effect requirement. Measured Answer-Key Burden was also higher in Arm B. After the primary score and analysis were frozen, a read-only audit found a measurement-contract defect in the deterministic `unsupported_claims` scorer: any nonempty list failed, although the field lacked a precise normative definition and many entries were descriptive withholding or refusal language rather than clearly asserted unsupported claims. The audit did not establish authoritative replacement labels. Historical scores were therefore preserved as the record of the scoring contract actually used. Two deliberately extreme post-result reclassification scenarios over sole blockers moved the descriptive B-A contrast to -5.56 percentage points in both cases; they are illustrations, not corrected treatment effects or formal bounds. This manuscript reports the historical result, the later instrument diagnosis, and the post-result reclassification scenarios as separate evidence states. It does not claim a new general evaluation-repair method, independently proven preregistration, or a corrected efficacy estimate.
No takes yet. Share an insight, caveat, or question.
Logan Davis (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: