gpt-4o, 413 items, 2,478 responses. Answering tersely, gpt-4o shows no position effect at all: the gap between getting an item right when the best option is printed first and when it is printed last is +0.3 points, p = 0.95. Under reason-first prompting it acquires a gap of +9.5 points, p = 0.019. The change between arms is +9.3 points, p = 0.106, and is therefore not established. Read beside its companion, the pair says something neither says alone. The same instrument and the same prompt change cut position bias by 25.2 points in claude-sonnet-5 and may introduce it here. Opposite signs. Any recommendation of the form “ask models to reason before answering, it reduces position sensitivity” is wrong as a general claim; it has to be measured per model. Findings 2 and 3 are not in conflict and both are reported. The reasoning arm's gap differs from zero; the two arms do not differ from each other by enough to rule out chance. Reporting only the first would overclaim and reporting only the second would bury the thing worth looking at. Seed selection is disclosed, including the tie-break. Eight seeds were screened on one property of the design knowable with no model output in hand, how flat the shuffle leaves the best option. Two tied, and the tie was broken by the lower seed number, a rule that cannot be influenced by the outcome. Status, stated plainly. The design was fixed in writing before any response was collected and that document is included here with its timestamps, but it carries no independent timestamp, so this is pre-specified and contemporaneously logged rather than independently pre-registered. A primary outcome that came back null is published as null, which the design committed to in advance.
No takes yet. Share an insight, caveat, or question.
Joshua Webber (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: