Accuracy-only evaluation rewards language models that guess over models that calibrate (Kalai et al., 2026). We measure its practical consequences in the bio-bibliography of classical Islamic scholarship, a closed domain where fabricated attributions are costly. We introduce ISNAD, a benchmark of 1986 verifiable Turkish claims across five probe types, among them contamination-resistant, evidence-conditioned anachronism probes. ISNAD also includes a 360-item bilingual (Turkish/Arabic) replication set with independently sourced gold. We evaluate ten primary models from six vendors under a matched no-reasoning configuration and three elicitation conditions (with supplementary arms and retained retries, 64,216 released raw records). Accuracy and reliability leaderboards rank different models first. We index reliability by committed accuracy, calibration, and selection-free discrimination. The raw-accuracy leaders rarely abstain and answer wrongly on 27–39% of all items. The strongest selection-free discriminator of real from fabricated attributions is a calibrated model (forced-choice d′ = 1.49 vs. ≤ 0.562 for the three raw-accuracy leaders) that answers 52% of items at 0.898 committed accuracy (expected calibration error 0.069). The leaderboards' scalar correlation is uninformative at ten models (Spearman ρ = −0.04, 95% CI [−0.66, 0.60]). A cost model (error plus λ·abstention) places the preference crossovers against the best-calibrated model at λ = 0.04–0.98. The dissociation survives forced choice, chain-of-thought, four prompt variants, a behavioral consistency channel, and ATTRIB_REAL gold-label noise to 35%. Blind expert studies validate the construct (expert-versus-gold κ = 0.725; fake-title AUC = 0.93). All probes, gold, and outputs are released, with evaluation recommendations for high-trust cultural-heritage settings.
No takes yet. Share an insight, caveat, or question.
Ali Çetinkaya (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: