Randomized trial evaluated outcomes of SSRA against a flat Transformer baseline, revealing negative results.
SSRA (Scale-Shared Recursive Attention) is a causal language-modeling block in which a single weight-shared rule per layer — one softmax attention block and one learned pooling operator, applied at every level of a binary tree over the sequence and reused by the token-level read-out — replaces the attention sublayer, at Θ(N·(w + m·log N)·d) training cost per layer: the same complexity class as Log-Linear Attention, and explicitly not a better one. A technical note (the stage-1 record of this work) fixed the design, its causality proof, its derived complexity, and a falsification plan before any training run; this paper executes that pre-registered plan, and the outcome is negative on every measured axis. The comparison is SSRA in its P1 (latent-query pooling) configuration against a flat Transformer baseline at matched parameters (≈ 84–85M) and matched tokens (850M; identical token stream, order, step count, and seed), with FLOPs and wall-clock reported, not matched. Parity: SSRA finishes +10.22 % above the baseline on validation perplexity at the training context length 1,024 (26.860 vs 24.369), outside the pre-registered ±5 % band. Stability: at the sweep-selected learning rate 1e-3 the SSRA arm suffered an unrecovered finite loss spike while the flat arm was clean on the identical token stream; the single permitted retune to 6e-4 — the only changed variable — trained cleanly in both arms, implicating the learning rate and leaving SSRA with an empirically narrower stable learning-rate range at this scale, mechanism undetermined. Length extrapolation (inference-only, N = 1,024 … 32,768): the flat baseline's pre-registered degradation prior is confirmed; SSRA's stable-to-mild prior is violated from N = 2,048; there is no crossover at any measured length. Retrieval (needle-lite passkey): SSRA scores 0 % in every length × depth cell, including its own training length, while the flat baseline shows depth-local copy behavior at the training length only; a pre-registered caveat records that an 85M-parameter model trained on 850M tokens need not exhibit copy behavior at all. Measured training-throughput constants on A100-class hardware favor the baseline by 11.8× at the reduced sweep scale (S1) and 11.1× in the production runs (S2), so the equal-token protocol favored SSRA on the compute axis. The repository releases the frozen implementation specification, the append-only decision log, the pre-registered assignments, per-run configurations, raw logs, and the verification suite. Either outcome was declared publishable in advance; this is the negative one. AI Assistance Disclosure This work used AI assistance throughout, under the author's direction. Design triage, formalization, drafting, and independent verification passes were assisted by Anthropic's Claude; the experiments were executed by Claude Code, an AI coding agent, on infrastructure operated and paid for by the author, under the author's supervision. The AI-executed work runs against machine-checkable guards: a frozen implementation specification as the normative reference for all code, one configuration committed before each launch, and a verification suite that includes causality (shift and completion), equivalence (frozen-reference A/B), and gradient-flow checks (§2.4, §3.6). All decisions, verdicts, and reviews are the author's. Consistent with the COPE position statement on authorship and AI tools, AI tools are not authors of this work; the author takes full responsibility for all content, including AI-generated portions. This section refines the record-level disclosure added to the stage-1 note's Zenodo record (DOI 10.5281/zenodo.20647034). Version history v1.0 (2026-07-19) — publication release (reviewer-pass gate PASS, 2026-07-18). DOI 10.5281/zenodo.21439493. v1.1 (2026-07-24) — DOI 10.5281/zenodo.21530947. Adds the weight-trajectory analysis of the §4.4 pair as §5.6 (pre-registered C-T1 criterion; verdict inconclusive; observations (a)–(c); flat control), figure/table upgrades — F5 pooling-entropy overlay (§5.1), F6 bf16 position-quantum figure (§5.3), Table T3 needle categorization (§5.5), F7a–F7c trajectory figures (§5.6) — and the cross-reference to the in-file erratum mirror of technical note v1.1 (§2.8); no v1.0 conclusion modified; one precision refinement (§5.2).
No takes yet. Share an insight, caveat, or question.
Daniel Šopov (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: