Handing the judgment work in a long-horizon task to a cheap, specialised judge is a divisionof labour usually sold on the grounds that it **saves money**. This paper measures that claimagainst a controlled three-way comparison, and reports four claims — one of which our owndata refutes. **First, cost is not the binding constraint.** We measured the per-call cost of threejudgment layers: a local non-autoregressive judge (Laya, marginal cost ≈ 0, excludinghardware amortisation), a remote typed-decision service (TypeSafe Jev, **$0.0000146–$0.0002406**,rising monotonically with input tokens, over 8 measured points on the direct route), and afrontier autoregressive model (DeepSeek-V4.1-Flash, $0.0000326–**$0.00009645**; at its highestreasoning setting $0.0001010, or 2.54× (**⚠️ eighth-round re-measurement: the maximum is 2.32x, and it falls at the LOW effort level, not max -- the multiple varies between draws**) the thinking-disabled setting). **At their respectivesmallest states** all three sit in the 10⁻⁵ dollar range (**the most expensive LLMconfiguration costs about 1.2 cents to complete one 120-checkpoint run**) — **but that orderof magnitude does not survive an increase in state size**: Jev costs **$0.0002406** for asingle call at its largest measured state (39,927 characters / 5,728 tokens), which projectsto about **2.9 cents** per 120 checkpoints; the LLM, extrapolated from its final state of1,751 tokens, projects to about **2.4–3.9 cents**. So "an order of magnitude cheaper" doesnot hold on any accounting, while "absolutely small" holds **only at short state sizes**.What actually separates the three is **latency** (p50 spans about **25–32×**: 37.4 ms /671 ms / **1.19 s** (one run of n=20; the pooled n=35 median is 1.07 s) — Jev measured as **independent wall clock**, a single run of n=20 at p50 1,192 ms, with thepooled n=35 median at 1,073 ms; **only one run has an artifact behind it, so no run-to-runrange can be reported**) and the **usablestate window** (local english clamps at 512 tokens, about 3,082 characters; the remote servicemeasured ≥15,002 characters — about 4.9× — **and that multiple is bounded by this harness'saccess layer: pushing the direct route to 39,927 characters still showed input tokens growingmonotonically with characters, which proves the "16,000-character cap" is the plugin's, notthe provider's**), and the **observability of failure** (the local one discards inputsilently, the remote one warns). We therefore restate our central claim from "cheaper" to"does not occupy a sequential round trip". **Second, and this is the paper's main contribution: these judges' self-reported fields arenot trustworthy, and their failures cluster in one place.** Seven independent phenomena werereproduced — **six from real calls**, the other from a synthetic backend driven purely by aninput hash (which produced a **completely credible results table**: Brier 0.359 on the 10 of14 battery items that carry binary ground truth, worse than a constant 0.5 predictor's 0.25).The six on the real-call side are: the truncation flag is wrong in both directions (englishlags by 111 characters; multilingual and typed-decisions false-positive early by 4,773 and3,774 characters), while the three checkpoints' flags all fire at the same point, 3,193characters, though their true clamps differ by 2×; `fits: true` and an already-truncatedinput can hold simultaneously; the probability field flips an item's conclusion on twodifferent systems (affecting half the items on one of them); and the two verdict words`conflicted` and `undecided` are unreachable under real inputs. These converge into a singleshape: **the judge is near-perfect when the answer is explicitly stated (0.9909, n=220), andcollapses when it must notice that something is absent or that something somewhere does notmatch (`no_support` 0.3091, n=220), and in both cases the confidence it self-reports is notlow.** The shape recurs across **three independent settings** (**silent truncation,`no_support`, cross-language** — the basis for grouping the three is in §7.5/§7.7; multi-hopchained verification is a **fourth** setting, but its Laya failure has a different shape, soit is not counted among the three); the sharpest of them is that when the candidate valuenever appears in the input at all, the judge gives **P(true)=0.5643** on average, i.e. itreads "not stated" as support (n=220). > **A correction that must be given alongside this**: we once described this judge as> "negative information". A Murphy decomposition of the n=1100 calibration corpus gives> **REL 0.0556 / RES 0.0397 / UNC 0.2400**, and **AUC = 0.7136 (95% CI [0.682, 0.745])** —> **resolution genuinely exists; the failure is in calibration, not in information.** The> correct description is **poorly calibrated**: its Brier (0.2571) is indeed worse than a> constant predictor (0.2400), and ECE is 0.2259. **Third, a negative result.** In the routing and cascade literature, using a non-generativediscriminator as a first layer has precedent (e.g. the BERT-style routers surveyed byRouteLLM), but to our knowledge no work has treated it as a **controlled first layer** andtested **paired complementarity** against a frontier generator. We ran that paired comparisonon **three** task regimes (the LLM arms are sampled through an API, so each arm reports**multiple independent draws** rather than a single point):- **Authority location**: the **forced-choice arm** has the LLM correct on all 48/48, the judge at 0.4583, and **it catches none of the items the LLM misses**; **⚠️ the prose arm of the same 48-item battery is the opposite**: LLM **46/48**, **1 judge-only item**, and **Δ_catch = +0.0435** (95% CI [−0.386, +0.471], resting on **2** items where the LLM errs). **Among the regime-level readings of the three regimes, this is the only positive Δ_catch point estimate**; the paper reports it as such, together with its width and its denominator (§8.2);- **77-class intent classification** (n=40, **4 draws**): LLM **0.750–0.900** (the four were 0.750 / 0.900 / 0.875 / 0.875; **median 0.875, mean 0.850**), judge **0.225** (identical across all 4 draws); **judge-only-correct 0–2 items against LLM-only-correct 23–27**, and **Δ_catch is negative in 4/4 draws** (−0.033 / −0.250 / −0.029 / −0.257) — the recorded run is the one **most favourable** to complementarity among them;- **Multi-hop chained verification** (re-judged after the ground-truth fix, n=68, **3 draws**): LLM **0.662–0.677**, judge **0.294**; **Δ_catch = −0.182 to −0.247, negative in 3/3 draws** — but **the interval strength is limited** (an unpaired Wald interval excludes zero in 2 of 3; under score/Newcombe **only 1 robustly excludes and 1 sits at the boundary**), and **failure correlation φ stratified by difficulty is not significant in any of the three** (CMH permutation p = 0.059 / 0.055 / 0.201; **⚠️ this group of p-values is not verifiable — no artifact, no script, no recorded seed, see §8.6.1(d)**). → **The conclusion is therefore stronger than "no complementarity found"**: the judge does notmerely fail to cover the generator's errors — its failures are **positively associated** withthe generator's. `P(judge correct | LLM wrong)` = 0.13–0.17, **below** its marginal accuracyof 0.294, while `P(judge correct | LLM right)` = 0.36–0.38 is **above** the marginal; thefailure correlation φ is positive in **3/3** draws (+0.19…+0.26).> **⚠️ Strength qualification**: the 2×2 Fisher exact p for that association is> **0.086 / 0.049 / 0.163** (r1/r2/r3) — **uncorrected, only 1 of 3 is significant at α=0.05;> under this paper's own pre-declared Holm rule, 0 of 3 survive**. And the 95% CI for the> one-sided test of "conditional accuracy − marginal" **contains 0 in 3/3 draws**. So this is> evidence **consistent with** a shared failure mode, not an established significant effect.> **Dissimilar does not automatically mean complementary, nor automatically independent.** > **Power and reproducibility must be given alongside the conclusion**: for the third regime> (after the ground-truth fix) the Δ_catch **MDE (80% power) = 0.28–0.30**, **still above the> pre-declared +0.10 gate** — so although the point estimate is **negative in 3/3 draws**,> **the interval strength is limited**: 2 of 3 exclude zero under an unpaired Wald interval,> **but under the zero-cell-safe Newcombe only 1 robustly excludes and 1 sits at the> boundary** (§8.3). **The precision remains limited.** The second regime's Δ_catch magnitude> is unstable (−0.029 to −0.257), **but its sign is negative in 4/4 draws**.> **⚠️ The sampling scope must be bounded**: **only two complementarity regimes** (the chained> battery `P22b` and the 77-class regime `P15b`) were re-sampled at `temperature=0` with a> reported per-item label agreement rate (chained regime **86.8–89.7%**; judge arm **100%**).> **The remaining LLM arms (P14 authority location, P19 calibration, P21 thinking cost, P23> logprobs, P24 horizon) remain single draws**, and their point estimates likewise should not> be read as precise values (§10.1). **Methodologically**, this paper reports four instances of the same class of error —**a specification-level semantic defect masquerading as a finding about the model** — eachfound by a control rather than by review, and derives **23** protocol mandates from them. **We state the boundaries of this paper explicitly**: every conclusion comes from measured**single-step judgments**; **autonomous long-horizon (multi-step cumulative) runs were notexecuted**, and the horizon slope is demoted to descriptive by pre-registration. All Layajudgment measurements use the **english checkpoint** at a **512-token** window (english loadedalone gives 1024, while multilingual/typed-decisions measure 1024 on the same machine) — sothe window is a function of **(launch loadout × queried checkpoint)**, and "window = 512" isnot a general property of the engine. On verification-style tasks the LLM hits the ceiling(correc
No takes yet. Share an insight, caveat, or question.
Perry Link (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: