Methodological framework demonstrates evidence admissibility rules after harness failures in multi-stage AI evaluation, highlighting valid claim boundaries post-repair.
Multi-stage AI evaluations can fail after some observations have already been produced. Provenance standards can reconstruct execution and artifact lineage; failure-localization and harness-repair methods can identify responsible components and repair surfaces; preregistration constrains post-outcome discretion; and claim-evidence systems can represent support, conflict, and missing evidence. A less explicit problem appears when these functions interact: after a failure is discovered and a repair is contemplated, what remains admissible from the historical evaluation record, and what maximum claim does that evidence still license? We present Evaluation Inference Custody (EIC) as a provisional evidence-state decision protocol for this problem, not as a replacement for provenance, reproducibility, failure localization, preregistration, or scoped repair. R3 replaces outcome-symmetry language with a predeclared dependency-closure carry procedure; expands experimental identity to distinguish declared design from implemented measurement generation, runtime, stochastic policy, and source snapshot; defines claim-relative completion witnesses outside the implicated failure closure; strengthens the conventional comparator; and requires an independent gold standard and explicit HOLD-cost accounting. A longitudinal governed-agent case illustrates why these distinctions matter while preserving null, mixed, unresolved, and adverse findings. The case remains hypothesis-generating. EIC should be retained as a distinct method only if preregistered prospective comparison shows reliable inferential decisions beyond strong conventional practice without offsetting evidence loss or abstention burden.
No takes yet. Share an insight, caveat, or question.
Logan Davis (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: