Experimental evaluation demonstrates reduced silent error rates in small specialist language models, indicating external deterministic verification surpasses intrinsic self-evaluation.
Multi-agent and small-model LLM systems commonly try to obtain reliability by asking a network to evaluate itself: self-reflection, verbalized confidence, or trained refusal. The research record shows that this pattern is fragile. We describe the Verification-Anchored Federation (VAF), a zero-trust operational harness in which honesty is a property of deterministic machinery outside the model weights rather than an emergent property of any model. Small LoRA-specialized adapters over a shared, CPU-friendly 1B backbone are wrapped in a deterministic router with out-of-distribution bounce-back, deterministic executors and tiered claim verification, and a per-specialist split-conformal abstention gate with drift-triggered recalibration. We report the co-primary metrics that the design forces to be read together: the silent-error rate (a wrong answer delivered with no abstention signal) and the abstention rate on solvable tasks. On a frozen 600-row adversarial benchmark across three domains, VAF reduces silent error from 36.17% (unverified 7B generalist) to 12.00% at first pass and to 4.00% (95% CI [2.70%, 5.88%]) after two mechanism changes, with abstention on solvable tasks falling from 38.67% to 24.00%. A fourth domain is added by the same pipeline in 9.0 wall-clock hours with no bespoke architecture. On four independent public benchmarks the frozen system attains near-zero silent error, but almost entirely by abstaining, and its calibrated conformal threshold does not transfer under distribution shift; on a 250-sample adversarial set authored by an unrelated model family it answers most traffic at 1.6% silent error versus 58.8% for the baseline. We report a corrections log of every published number we retracted, and argue that the negative results are as load-bearing as the positive ones. Preprint, not peer reviewed. 14 pages, 1 figure, 7 tables.
No takes yet. Share an insight, caveat, or question.
Amol Pol (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: