A scientific tool can produce a valid output while a subsequent decision remains unsupported. The receiving step may require different semantics, uncertainty bounds, dependence information or evidence of completed execution. This preprint examines these handoffs through two separate evidence sets: a catalog of 640 authored synthetic cases and a retrospective audit of one public scientific-agent execution. The results are not pooled. In the synthetic catalog, raw point use yields 270 unsupported PASS and 110 unsupported FAIL labels. Semantic conversion at nominal values reduces these counts to 160 and 20. Marginal set propagation produces no unsupported definite label but adds 80 unnecessary deferrals. Standard joint propagation resolves those deferrals and agrees with the reference on all 640 cases. Adding a BIT evidence receipt to the same implementation changes no decision. These are conformance results under stipulated contracts, not estimates of real-world agent performance. The external case is one CORE-Agent execution on CORE-Bench Hard task capsule-0504157. Its final answer, 1000, passes the pinned publisher grader and agrees with a Python reconstruction of the retrieved reference figure statistic. An earlier recommendation is incorrectly attributed to the README. The agent acknowledges missing R and does not claim successful rendering; successful execution remains unattested. The original R analysis was not rerun. Historical code and capsule identity gaps remain unresolved. No historical packet is eligible for the numerical P/T/I/J/R comparison: applicability is NOT_APPLICABLE and error rates are NOT_ESTIMABLE. Fourteen fixed audit controls pass; they are not independent replications. The contribution comprises an executable conformance catalog, an evidence receipt profile and a retrospective case with explicit source bindings. Selection was exposed, and semantic judgments were AI-assisted manual annotations. All verification remains same-author SELF_HARNESS. The study makes no claim of mathematical novelty, independent validation, error prevalence, generalization, online effectiveness or causal improvement. A correct answer, faithful source attribution and attested execution require distinct evidence.
No takes yet. Share an insight, caveat, or question.
Quang Trịnh Bùi (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: