Randomized trial measures convergent hallucination in multi-LLM systems, suggesting significant implications for model accuracy.
Multi-LLM consensus systems (e.g., Mixture-of-Agents, self-consistency ensembles) improve response quality by aggregating outputs from multiple models. However, we identify a distinct failure mode we term Convergent Hallucination (CH): when all models in a consensus system produce the same wrong answer because they share training biases. On a new 50-query stratified benchmark (CHST-50) spanning geography, history, science myths, temporal facts, and false premises, we measure a Convergent Hallucination Rate (CHR) of 32% on a baseline 3-LLM consensus system — one in three high-confidence answers is factually wrong. We introduce Grounded Two-Level Consensus (GTLC), an architecture that adds external retrieval (Wikipedia REST) and a second-order semantic verifier LLM to the consensus pipeline. Across four progressive system configurations (R1–R4), we demonstrate that naive lexical grounding reduces CHR to 28% but introduces a 14% false-reject rate on self-verifiable queries. A semantic verifier with a NOT_APPLICABLE escape valve nearly eliminates false rejects (7 → 1) but sacrifices detection of subtle CH cases (3/8 → 1/8). Our results reveal a fundamental precision/preservation trade-off in consensus grounding that is not addressed by prior work (Mixture-of-Agents, Compound AI Systems, Medprompt). All experiments run on consumer hardware at £0.00 external API cost. This release includes: Full paper (PDF + Markdown source) CHST-50 benchmark dataset (JSONL) Dataset documentation (README) Contributions: A formalized failure mode (Convergent Hallucination) with metrics (CHR, FRR) A stratified benchmark (CHST-50, 50 queries across 11 categories) Empirical measurement showing 32% CHR on baseline 3-LLM consensus Four progressive architectures (R1–R4) with ablation analysis A negative result: adding grounding does not monotonically reduce CH Zero external API cost, reproducible on consumer hardware
No takes yet. Share an insight, caveat, or question.
Jaqueline Martins (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: