Empirical analysis reveals monotonic scaling in knowledge application alongside early saturation of cloze retrieval across language models, indicating capacity-dependent architectural bottlenecks.
Large language models solve two differentiable task types on the same underlying knowledge base. Cloze retrieval saturates early (~90 % by 8 B parameters); application scales monotonically across three orders of magnitude. We document this asymmetry on a single invented knowledge domain ("Zorbetik") across nine models, using frontloaded in-context learning (Brown et al. 2020). Four findings: (1) application scales monotonically with capacity (Spearman ρ = +1.000 on Qwen2.5; cross-family panel ρ = +0.92, n=9; +40.8 pp per decade); (2) the bottleneck migrates with capacity; (3) Mixture-of-Experts models scale on active parameters, not total; (4) at uniform n = 47 with Wilson 95 % CIs, four capable instruction-tuned substrates cluster in a capacity-independent 79-85 % application band, the empirical signature of a race-architecture floor (proposed in Paper 10 §1.5). A companion paper (Paper 2C, in preparation) re-tests the floor on an independent 327-substance domain and on the chain-depth axis. Companion papers in the series: Paper 0 (BFT): 10.5281/zenodo.19462500 Paper 1 (per-token measurement standard): 10.5281/zenodo.20012654 Paper 2B (substrate-mechanism companion, in preparation) Paper 2C (chain-depth axis companion, in preparation) Paper 3 (Friction-guided inference): 10.5281/zenodo.20014122 Paper 10 (Race-architecture, physics scope): 10.5281/zenodo.20014568 Data, fine-tuning notebooks, analysis scripts: https://github.com/tplund/friction-theory-p2-capacity-scaling v6 (August 2026) — attribution and scope revision. The shape of Finding 1 — single-fact retrieval improving with scale faster than multi-step composition does — is the compositionality gap, measured and named by Press et al. (2023) in the GPT-3 family. This version says so where the finding is stated, and keeps the increment it earns: the gap is shown here on a domain no model can have been pretrained on, so it cannot be a residue of uneven pretraining exposure across single-hop and multi-hop items. Finding 2's plateau is placed inside Dziri et al. (2023) on error compounding with compositional depth; the open question is what sets its height, not that it exists. The invented-domain design is credited to the synthetic-world tradition (PrOntoQA; ProofWriter) rather than claimed as new, and the scorer-disagreement result is placed inside Kamalloo et al. (2023). The word “capacity” is now credited to the tradition it is borrowed from (Miller; Baddeley & Hitch; Cowan; Just & Carpenter), with an explicit statement that no claim is made that a language model’s in-context bound is that construct. In the other direction, the audit’s cross-family blind step is now reported at its own size: key agreement of 35/35 on the valid items is the strong result, while the re-adjudication at κ = 0.628 is a sensitivity analysis by one language model of another, with no human agreement collected. A companion study previously cited as forthcoming has been parked, and the text no longer promises it. Editorial pass. Earlier versions remain in the version history.
No takes yet. Share an insight, caveat, or question.
Tomas Pødenphant Lund (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: