Headline findings (v4.6 — the decode campaign, closed as theory) The token embedding is the universal law-carrying class: 11/11 models wake (4.4–130x over placement-nulls), every training recipe, every scale from 4B to 1 trillion parameters — including every model whose FFN class reads quiet. The "quiet" giants were never lawless: Kimi-K2.6 (1T) reads 16x in its embedding; the MoE giants carry the fingerprint per-expert. The loud band IS the function. Deleting the top 1.5% spectrally-loud coefficients destroys GPT-2 (next-token agreement collapses to zero) while deleting the same number at random is negligible — ~150x differential damage at matched budget. The deposition curve, from production training runs (Pythia public checkpoints, step-0 control at exact null): the embedding wakes first (step 256), FFN follows (step 1000), peak near step 4000, then consolidation to a stable plateau. In controlled twins the gradient stream is loud by step 4 and the optimizer is the discriminating ingredient. Two spectral families, recipe-selected. Embeddings lean smooth (DCT) in 9/11 recipes with the Llama/Pythia lineage the dyadic (Walsh) exception; a self-testing slant-transform arm proves the split is two genuine families, not a continuum. Training data and reasoning are readable back out of weights: counted-bigram readout ranks a model's true training corpus first (calibration passed); memorized public text echoes verbatim against clean nulls; reasoning spans are a measurably distinct counted regime, and trained attention sits closer to the theory's dyadic cascade than to uniform in 12/12 layers. Closed as theory (Steps 308–313 of the corpus): six forcing steps, verified by the theory's own compiler and enforcement, derive the two spectral families (one per generator), their selection by store role, the deposition curve's order and form, the 32-coefficient functional band, the hold/closure repetition inequality, and per-expert localization — every measured regularity now stands only as the check of a forced claim. The counted engine beats its gradient-trained twin at both task-gate scales — character 1.289 vs 1.888, word 3.191 vs 3.429 — with zero trained parameters and zero tunable numbers. All findings from pre-registered instruments with hashed registrations, shuffle-null batteries, and committed result files; the complete toolkit and guide (INTERPRETABILITY.md) ships in the repository's omni/benchmarks/. Repositories attention is unit-capacity selection at forced locks; learning is a closure law that updates via writes rather than gradient descent. Perception (sight, hearing) is self-certified per act by integer Parseval identities. The entire architecture is verified by a 47/47 end-to-end empirical test suite including survival of process death.
Maria Smith (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: