Randomized trial evaluates benchmarking outcomes in emergency department injury data, highlighting implications for AI models.
An external, deterministic, model-free grounding audit of the Rafe and Das (2026) helmet-abstention benchmark (1,464 emergency-department injury narratives). A model-free gate re-reads each note against a sealed schema and returns a three-valued verdict; it reproduces the benchmark's separation between fabricating and abstaining models with no model in the loop, exposes nine likely gold-coding slips, and shows that a single change of prompt wording moves even a frontier model into fabrication. Version 2 corrects version 1, which contained the report only and described a harness, data, and per-cell outputs as released that were not in fact included. This version includes the evaluation harness, the Rafe and Das data as used (CC BY 4.0), and the complete per-cell outputs, and adds the adversarial negation-and-scope study the report named as its natural next step: the sealed schema's false-grounded rate under adversarial cancellation is 27 to 34 percent (development and held-out), reduced to about 5 percent by a candidate scope-hardened schema (not yet sealed), benchmarked against NegEx, ConText, and a natural-language-inference baseline; the benign-corpus false-grounded rate is zero at an effective sample of about fourteen (Wilson upper bound near 20 percent), and the nine flagged cases are gold-coding divergences. Withheld under UK patent application GB2606072.3, pending the doctoral thesis: the OFM-PBL core (the verdict-folding engine) and the sealed schema. The harness scripts import them and are provided for inspection; every reported figure is verifiable against the released per-cell outputs, but verdicts cannot be re-derived from raw notes without the withheld core and schema. See the Version 2 note in the report, DEPOSIT_README.md, and DATA.md.
No takes yet. Share an insight, caveat, or question.
David Antonio Tomé (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: