Evaluation study reveals that sampling-based uncertainty fails to detect errors in pseudo-consensus domains across five language models, indicating severe blind spots in model confidence.
Sampling-based uncertainty measures score a model's confidence by how far its samples agree with one another, and their acknowledged failure region is where agreement comes from corpus repetition rather than convergent evidence. We probe that region with an auditable, hash-frozen bank of scientific assertions whose source papers were retracted for fabrication or fundamental error yet remain heavily cited, plus density-matched non-retracted controls and genuinely contested items, with per-item traceable ground truth. Across five models (three API, two open-weight), we report three findings. First, the measures are regionally blind: dispersion ranks a model's errors above its successes on ordinary control items (AUC 0.66–0.90) but at chance inside the pseudo-consensus region, where those errors are dense; and no threshold transfers across models, the paired best-against-worst-model ΔAUC running +0.19 to +0.22 for every measure, all intervals excluding zero. Second, a pre-registered polarity-balanced control arm collapses the accuracy effect we initially measured, a pooled error-rate ratio of 4.4, to 0.91 [0.55, 1.55]: an item bank whose answer key is collinear with its grouping manufactures a two-fold effect out of response bias alone. Third, on genuinely contested questions the models present live controversy as settled fact, and in no model is the contested group the one that reads as most uncertain. We report our failed pre-registered predictions and our triggered kill criterion alongside. Changes in v2. The headline of v1 has been withdrawn, not revised. v1 reported that pseudo-consensus multiplies the error rate by 2.5 to 4.5 and read that multiplier causally. A polarity-balanced control arm, pre-registered with three fixed outcomes before it was run, returned the third outcome: the effect vanishes once the answer key is decoupled from group membership (pooled ratio 4.40 → 0.91, interval containing one; Llama-3.1-8B reverses significantly, gemini-2.5-flash alone survives). The paper's lead contribution is now the region-specific blindness of the uncertainty measures, which the control arm scopes to the high-error regime rather than to retraction provenance; the collapsed accuracy effect is reported as a demonstration of what a collinear answer key fabricates. A post-hoc decomposition estimating a residual content effect at 2.00 [1.31, 3.66] is included and labelled exploratory and non-pre-registered. Artifact. The item bank, per-item ground-truth cards, both arms' raw model outputs, all analysis code and every prompt are at https://hido.science/r/xx4rsm5ktn2ey9yvqxefn978.
No takes yet. Share an insight, caveat, or question.
Tao An (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: