Retrospective audit reveals rubric-specific score gains without decision-quality improvements across AI prompt augmentations, highlighting potential criterion contamination.
This paper presents a retrospective exact-cell audit of an interrupted and sequentially hardened prompt-augmentation experiment. The audited record contains 306 evaluated generations: 252 generations from seven domains completed in the initial tranche, plus a separate 54-generation check in three domains selected after the interim and validity sequence. The two strata have different selection and generation histories and are never pooled. The source inventory also contains 15 unevaluated F2 ablation generations and 18 legacy generation files without draw identifiers; neither group enters an analytic denominator. The initial outputs were scored with a construct-named `/35` rubric recorded as explicitly rewarding concepts supplied by the treatment, and later with a `/10` LLM-judged decision-quality instrument and `/2` LLM-judged correctness instrument. Under a frozen paired estimator—`B-A` within `(domain, subject label, draw)` after averaging the two recorded judge values—42 paired cells use the A/B rows from 84 of the 252 initial generations. Their construct-named contrast is `+9.25`, with a descriptive whole-domain resampling range of `8.15` to `10.35`. The LLM-judged decision-quality score on the same outputs is `-0.20` (`-0.39` to `+0.01`), and the LLM-judged correctness score is `-0.04` (`-0.17` to `+0.08`). The construct total and LLM-judged decision-quality score have Pearson/Spearman associations of `0.107167/0.083923` over all 252 initial generations and `0.092225/0.072939` over the 126 `A/B/C` generations. In the post-selected hard stratum, 18 paired cells use the A/B rows from 36 of 54 generations; LLM-judged decision-quality `B-A` is exactly `1/36`, displayed as `+0.03`, while LLM-judged correctness `B-A` is exactly zero. A retrospective audit of the historical H6 joint selector-and-arm comparison—`haiku,B` minus `sonnet,A`—yields `+8.64` on the construct-named rubric, but `-0.64` on LLM-judged decision quality and `-0.14` on LLM-judged correctness. Chronology is load-bearing. All three hard domains were selected and evaluated after the initial validity sequence. The 18 F2 `A/B/C` generations already existed before that selection, whereas the 36 P1/S1 `A/B/C` generations were produced later. Only two subject-model judges completed. Their labels share a product-family nomenclature with the subject labels, while exact model build identifiers and generation settings are absent. The strongest supported interpretation is that the evaluated cells exhibit an instrument-sensitive pattern consistent with criterion contamination. The study does not identify that mechanism, establish equivalence, or support a universal augmentation null.
No takes yet. Share an insight, caveat, or question.
Thomas Schwarzenböck (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: