A qualitative elicitation study in which three frontier large language models (SuperGrok 4.6 Expert, Claude Opus 5.5 High, GPT-5.6) answered the same seven questions about whether progress in frontier LLMs could become decoupled across capability, benchmark performance, reliability, real-world usefulness and properties relevant to robust general intelligence. The purpose is hypothesis discovery, not confirmation: agreement among language models is not evidence of truth. The paper separates model-generated hypotheses, cross-model recurrence, independently verified evidence, mechanisms, causal claims and falsifiable predictions, and checks every model-supplied citation or figure against primary sources. Five decompositions recurred in all three models, including capability ceiling vs. reliability floor and model-level vs. system-level capability. The strongest defensible proposition is that frontier capability gains do not logically entail proportional gains in robust general competence; whether they fail to do so empirically is open. The paper states ten hypotheses with counterarguments and proposes nine falsification experiments. A second round adds a cross-review, memory-off reruns and a control question. Anonymisation did not hold: reviewers recognised their own responses, except a memory-off reviewer (0/7). Several model-specific themes depended on the researcher's accumulated context (e.g. evaluation awareness: 26 mentions with memory, 0 without). The paper documents the full provenance of the hypothesis, including prior joint brainstorming between the researcher and the models published on X (@YaffFesh), and argues that in-context and clean runs are complementary arms of one instrument. Files: main paper (PDF); Supplement S1 (first-round transcripts); Supplement S2 (second round: cross-review, reruns, control, X-post context); LaTeX source with transcripts and analysis data. Author profiles: X https://x.com/YaffFesh · Bluesky https://bsky.app/profile/yafffesh.bsky.social
No takes yet. Share an insight, caveat, or question.
Yanush Feshter (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: