Large language models (LLMs) are entering decision-support systems as inexpensive synthetic experts, yet their judgments are rarely validated against published human targets. We benchmark six LLMs on a source-audited corpus of 27 published Analytic Hierarchy Process (AHP) studies with a preregistered holdout of 22 studies, 21 reference-weight tasks, 28,783 valid benchmark calls, and study-clustered inference. Holdout rank replication was modest and model dependent. A diagnostic showed that fixed-source-order rank evaluation is confounded with criterion presentation order: reference vectors largely follow source listing order, and a model-free descending-order rule reaches a mean Spearman correlation of 0.735, outscoring every model. A prospective randomized-order experiment raised Claude Sonnet 5’s mean Spearman correlation by 0.46, and a registered follow-up with an exactly balanced schedule replicated that gain. For HCX-007, the model that tracks presentation order most strongly, the randomized-order decline was directional but not statistically significant (0.17, Holm p=0.16), and the follow-up was inconclusive about the average difference (+0.007, 95% CI [−0.23, 0.22]); its single draws continued to follow presentation order, and in an exploratory subgroup of six tasks with seven or more criteria its balanced-order accuracy was lower. Most holdout panels could be implemented only as repeated draws under a shared profile, and matched individual human data were limited to two tasks. Benchmarks for synthetic experts should therefore balance presentation order and evaluate aggregate reconstruction, individual human-likeness, and downstream ranking separately.
No takes yet. Share an insight, caveat, or question.
Kim et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: