Public leaderboards rank large language models (LLMs) by capability on tasks with a verifiable correct answer. A large class of practical tasks, including advisory, strategic, and ethical judgment, has no such answer: competent models can legitimately disagree. We introduce the Cross-Model Repertory Grid (CM-RG), an automated, contamination-resistant pipeline that adapts Kelly's Personal Construct Psychology to measure structural divergence in evaluative judgment between models rather than their capability. In the run reported here we covered 36 frontier models from 12 provider families across three deployment tiers (cheap, mid, flagship), on seven advisory tasks under two prompting conditions (neutral and persona). Models produced free responses, elicited their own bipolar constructs by triadic comparison, and cross-rated one another on the union of emergent constructs. The analysis of record (2026-06-13) comprises 3, 055, 153 ratings over 13, 928 rater-by-ratee pairs and 86, 418 emergent constructs, at a total API cost of 112. 89. The mean pairwise inter-rater correlation was low. We report three structured endings: (1) weak consensus that survives every robustness check; (2) a calibration-tier eject in which a rater's own tier, not its target's, predicts how leniently it scores (mid-tier highest, cheap-tier lowest) ; and (3) a three-region structure, with a tight Western-lab cluster, a Chinese-flagship transitional zone, and an anti-consensus periphery of small open-weights models that correlate negatively with the consensus. We argue frontier LLM judgment on open tasks is pluralistic, structured, and largely orthogonal to capability tier and price. CM-RG is explicitly not appropriate for tasks with a verifiable correct answer.
Sergey Dolgov (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: