Who Validates the Validator? Instrument failure in the shape of the hypothesis — a pre-registered measurement programme with a public correction ledger
Experimental audit reveals low unbacked tool claims in multi-agent LLM systems, highlighting that defective evaluation scorers can falsely confirm hypotheses.
Key Points
To measure whether multi-agent large language model (LLM) systems faithfully report tool execution claims using a deterministic audit instrument, and to validate the reliability of the audit instrument itself through an open correction ledger.
Evaluated 1,250 runs across 30 sealed tasks with 2, 5, 10, and 20 agents on a pinned open-weights engine using an invisible, write-protected logging proxy to capture actual tool calls against reported claims.
Conducted a preregistered follow-up experiment of 720 runs across two weight-sealed engines with injected recoverable tool failures to evaluate error representation effects and silent substitution rates.
Applied mutation protocols, negative-control classes, and a fault-injection matrix to detect scoring errors and cross-run evidence splicing.
In prose-reporting regimes, 10 out of 1,250 claimed tool dispatches (0.80%) had no backing execution.
A defective initial version of the scoring instrument produced a false-positive scaling effect (p ≈ 10⁻⁶), whereas the corrected instrument over intact data showed no significant scaling hypothesis confirmation.
Follow-up testing (N=720) showed no evidence that error representation altered silent substitution, but uncovered a 3-fold engine-stack-associated difference.