Randomized trial evaluates a mechanism for skill compounding in AI systems, suggesting efficiency in training signals.
We tested the most-cited mechanism for compounding without retraining — the ever-growing skill library — in a compositional visual-reasoning domain, with held-out, exact-match verification at every step. It barely compounded: 4.0% of mined problems (95% CI 1.8–8.5%) produced a bankable, reusable program, because 88% of measured solutions needed exactly one novel step, and none needed more than two. What improved with experience was a lightweight learned ranking policy: it converts roughly 75% of experience into training signal (CI 68–82%) versus 4% for the library — a ~19× asymmetry (CI 10–58×) — and it made multi-step search tractable where blind enumeration had been silently skipped on 95% of problems. In a controlled synthetic family where solution depth is dialled directly, the library's gain is zero at depths 1–2, rises to +16–20 points at depths 3–4, and pays off only in a predictable window: between what blind search already reaches and what macro-extended search can reach. Under a starved budget the gain even turns negative — every banked entry costs branching. The general claim this supports: the depth distribution of a domain's solutions, relative to the searcher's blind reach, predicts where compounding will occur. The reporting discipline that holds it together: publish the slope of verified reuse, never the count of stored skills. Article 3 of a four-part series on measurement and honest growth in AI systems. The controlled depth-dial experiment is reproducible from the companion code: https://doi.org/10.5281/zenodo.21200069
No takes yet. Share an insight, caveat, or question.
Matthew Childs (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: