Experimental evaluation demonstrates that hybrid multi-tier memory boosts long-context reasoning and cuts catastrophic forgetting in language models, highlighting complementary learning systems.
Frontier models approach long-context reasoning by scaling the transformer context window: one of the several memory systems the brain evolved, scaled alone. We present Sapience, a hybrid architecture that builds the missing ones: a substitutable transformer cortex coupled to hippocampal episodic retrieval (L1), cortical-replay consolidation (L2), and sleep-stage weight adaptation (L3) — the complementary systems that Complementary Learning Systems theory says a complete brain pairs with a cortex. Four results carry the claim. (1) The architecture carries the long-context selection and scaling load; the reader sets much of the absolute accuracy level: holding reader weights fixed, a frontier reader answers 3-hop questions at 34.4% from a raw ~959K-token window and at 66.7% from ~700 tokens retrieved by the store (n=90 item-paired; McNemar p=1.5e-5; judge-invariant across five judges from three vendors). (2) The system operates at 10M-token history scale where raw-window configurations fail: on one continuous benchmark the reader's working set stays bounded while history grows 2,500-fold (flat ~82% at 2M-10M, 3 seeds); on a paired forward-growth construction the same reader scores 0/20 with the answer verified inside its delivered window while Sapience answers 45.0% on identical items (b=9, c=0, p=3.9e-3), at over 1000x lower cost per query; and on knowledge that changes, supersession structure beats relevance retrieval by +40.4pp on real Wikipedia revision histories — retrieval saw both values and could not tell which was current. (3) The consolidation tier has one measured system-level function, and a broader one still open: the write-time aggregation gate materializes exact distinct-counts at write time and beats a compute-matched LLM counter emitting the identical sentence by +22.8pp (95% CI [+8.1, +38.2]) on covered aggregation questions, attributable to deterministic count correctness; the same contrast runs -3.3pp against the gate on trend questions and is exactly neutral on single-episode controls; broader semantic and schema consolidation is implemented but unestablished. (4) Replay drawn from the same store makes adaptation durable: in an internally pre-registered matched-acquisition design (L3), replay generated from the episodic store cuts catastrophic forgetting by 47.0pp (95% CI [39.7, 54.0]; acquisition formally non-inferior), generalizes across model families, domain pairs, and a real git-history corpus (75.5/51.8/+52.9pp), outperforms a tuned regularization baseline, and survives dose- and step-matched ablations that isolate the replayed content, consistent with the CLS prediction that interleaved replay reduces interference. Under genuine conflicting updates, generic replay wins retention and composite current-world accuracy while paying a measured stale-value tax; restricting replay to the memories the store marks current removed every observed stale response (0/81 superseded probes; 95% upper bound 12.5% at n=27) at non-inferior retention and the programme's best composite current-world accuracy (90.7%); a volume-matched exclusion control does not reproduce the effect (13.6% tax, -7.9pp retention), so it follows from which content is excluded, not from replaying less. The architecture also predicts where memory should not help (injection where the cortex suffices costs accuracy; within the window a strong reader ties retrieval), and both predictions are confirmed by measurement. Every claim is reported with its scope: the primary mechanism, transfer, real-corpus, replay-format and conflict instruments are <=8B adaptation instruments, while a fixed-load scale ladder and a three-seed learning-rate-matched experiment extend measurement to 72B under a stated protocol asymmetry and a load-to-capacity confound; per-layer accounting separating measured positives from open bets, and an adversarial audit protocol behind every headline number (the headline table). Throughout, cortex names the architectural slot holding a standard LLM and reader that same LLM in its answering role; in every matched comparison both arms use the identical reader, and only what it reads changes.
No takes yet. Share an insight, caveat, or question.
Sapience Labs (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: