Randomized trial evaluates a two-layer guard reducing fabrication rates in regulated domains, indicating necessity for oversight.
Multi-LLM consensus systems are increasingly deployed in regulated professional services (immigration advisory, financial compliance, medical, legal), where practitioners hope that agreement among independent models will reduce hallucination risk. We document the opposite failure: when sub-models share overlapping training corpora with stale snapshots of a regulated domain, they converge on the same fabrication, and semantic-agreement scoring correctly reports high confidence in the wrong answer. On 2026-06-28, six sub-models produced 5-out-of-5 fabricated claims about UK immigration adviser rules at 78% consensus confidence. We introduce a two-layer safety guard: (1) a domain-aware regex heuristic flagging fabrication shapes (SOC codes, precise monetary figures, invented citations, regulator renames) at zero LLM cost, and (2) an LLM-as-judge second line that, when the regex layer flags, calls a cheap model to identify the shared prior the sub-models are converging on and returns a calibration multiplier. The layers are fused multiplicatively. We construct RegBench, a piloto benchmark of 10 claims across UK immigration and EU AI Act domains, each with ground-truth answers verified against primary authoritative sources (gov.uk, artificialintelligenceact.eu) on 2026-07-11. On this piloto, a single-model baseline (DeepSeek Chat) fabricates in 80% of prompts. Our two-layer guard reduces neither the fabrication rate nor the answer content — it flags 75% of fabrications for downstream caution at a marginal cost of £0.00005/query. We report this as calibrated risk classification, not correction. The guard is particularly relevant to Article 14 (human oversight) obligations under the EU AI Act, which take effect 2026-08-02 for high-risk systems. Code, benchmark, and reproducer are released under FSL-1.1-Apache-2.0.
No takes yet. Share an insight, caveat, or question.
Jaqueline Martins (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: