claude-sonnet-5, 273 items, 1,638 responses. On a clinical research competency instrument, the gap between answering correctly when the best option is printed first and when it is printed last falls from 35.3 points to 10.0 points when the model is asked to think before it commits. Difference-in-differences −25.2 points, z = −4.40, p < 0.0001. The mechanism is visible and it is not a capability gain. Reasoning barely moves position 4, from 87.3% to 85.8%. It lifts position 1 from 52.0% to 75.7%. The model does not get better; it stops losing the answer when the right one comes first. Anything that reads this as a general accuracy improvement misreads it. The shuffle null was verified before the result was read. At the pre-specified seed the best option sits at 25.5 / 24.4 / 24.4 / 25.8 percent across the four positions, n = 792, largest deviation 0.8 points. The effects measured are 10 to 35 points. Every figure here reproduces from the retained rows. The answer key is regenerated deterministically at the run's own seed, so the analysis is a re-scoring of stored responses rather than a second implementation. A 100-item pilot of the same contrast measured −14.1 and was published as not established; the powered study finds −25.2. An underpowered pilot is not a small version of the answer, it is a noisy one. Status, stated plainly. The design was fixed in writing before any response was collected, and that document is included here with its timestamps. It was written to a session scratchpad seconds before the run launched and was never independently timestamped, so this study is pre-specified and contemporaneously logged, not independently pre-registered. The failures are in the record too: the run was destroyed twice by an unretried timeout, four earlier runs were voided, and the void run ids are named rather than dropped.
No takes yet. Share an insight, caveat, or question.
Joshua Webber (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: