The Judgment Layer Turning agent memory into wisdom — and the layer nobody has built Author: Mike Norton · ORCID 0009-0003-1866-6249 · DHARE — dhare.com.au · Brisbane, Australia Version: 1.0 · August 2026 · License: CC BY 4.0 Status: Architecture and evaluation protocol. A results paper follows the build. The judgment prompts, gate thresholds and evaluation corpus stay internal. Keywords: agent memory · judgment layer · cognitive architecture · constitutional governance · continuity They built the library and the catalogue. We’re building the librarian who decides what the library is for. Abstract Persistent memory for AI agents is becoming a commodity. The funded field — Mem0, Zep/Graphiti, Letta, Cognee, LangMem, and now platform offerings like Cloudflare’s Agent Memory — handles extraction and provenance well. But agent memory breaks down into four steps: extract what happened, attribute where it came from, judge what it means, and govern how it changes behaviour — and every shipping system stops after step two. Recent formal work (Roynard, 2026) identifies the same missing tier and reports near-zero contradiction-resolution across the field. I stake four claims about that empty rung. One: memory should be shaped like attention — a tree of focus sessions with depth-proportional compression — not shredded into atomic facts. Two: continuity belongs to the judgment layer, not the model — treat the LLM as stateless by design and identity survives model swaps, context compaction and provider churn. Three: typed persistence needs an ignorance ledger — a first-class store for known unknowns, with suspension of judgment (epochē) as a routing outcome alongside accept and refuse. Four: the layer that turns memory into behavioural guidance has to be constitutionally governed — promotion gated on evidence rather than user approval, amendment restricted to the human owner, and no agent ever self-ratifying the law that governs it. I finish with an evaluation protocol — architectural property tests plus a continuity test I call the ache test — and an open invitation to tear it apart. 1. Everyone files. Nobody judges. Here’s the state of play. The current generation of agent-memory systems is genuinely good at remembering. Mem0 extracts atomic facts at scale. Zep’s Graphiti tracks temporal validity with bi-temporal timestamps. Letta treats context as RAM and lets the agent page its own memory. Real achievements, and this paper builds on them, not against them. But remembering is the easy half. Break agent memory into its four steps — extract, attribute, judge, govern — and you find that steps three and four “require custom work” in every system surveyed, and no system natively implements step four at all. Roynard (2026) formalises the missing tier: a four-layer model where Knowledge updates by supersession, Memory decays unless consolidated, Wisdom updates only through evidence-gated revision, and Intelligence is ephemeral inference. His benchmark discussion reports near-zero contradiction-resolution scores across current systems. And his pilot shows this cuts both ways: typed routing beat a flat memory store by +0.128 overall (+0.106 on contradictions, +0.150 on temporal reasoning) — while a naive keyword router reversed the entire advantage (−0.125). Judgment isn’t a garnish on memory. Mis-routed memory is worse than no memory. It gets worse before it gets better: the benchmarks that should referee this are themselves broken. A 2026 community audit found LoCoMo’s answer key about 6% wrong, its LLM judge accepting most wrong answers, and LongMemEval fitting inside a single modern context window (both as reported in Roynard, 2026). The field is optimising scores that don’t measure the thing that matters. So the open problem isn’t storage, retrieval or extraction. It’s judgment: what deserves keeping, what earns the right to direct behaviour, how contradictions resolve, what gets forgotten, and who governs the whole process. This paper is an architecture for that tier. I call the layer Marcus, and Section 5 explains why that’s not just a name. 2. Claim one — memory shaped like attention Every system above shares the same reflex: shred experience into retrievable units. Facts, triples, summaries, embeddings. The shape of the thinking gets thrown away and only the residue is kept. I think the shape is the point. Human episodic memory is organised by attention, not by fact-type: a subject held in focus for a duration, branching into sub-topics, re-entered later at the depth you left it. So the unit of memory in this architecture is the focus session — a tree where each node is a subject of attention with a natural depth class (a ten-minute glance, an hour’s focus, a six-hour deep dive) and the children are the branches attention actually took. Two things fall out of that. Compression is proportional to depth. At session close, a glance compresses to a line or to nothing; a focus to a paragraph with episode references; a deep dive to a structured summary with its branch tree intact. Re-entry is a first-class operation. “Where were we on the memory-architecture deep dive?” resolves to a node, restored at stored depth, with its open branches — not to a similarity-ranked pile of fragments. And there’s a nineteen-century-old existence proof that this compression discipline keeps what matters. Marcus Aurelius’s Meditations is the output of a nightly consolidation practice run for about a decade by the bloke administering an empire — and it contains no transcripts. Thousands of audiences, dispatches, campaigns and crises compress into twelve thin books of lessons, some entries stamped by location (“Among the Quadi on the River Gran”). The trivial day left no entry. The repeated lesson became character. Ten years of raw experience down to a portable wisdom artifact that’s still governing readers — that compression ratio is the target function for a forgetting mechanism, and it says aggressive, judgment-driven discard loses far less than intuition suggests (consistent with the roughly-20%-of-facts-preserves-full-behavioural-fidelity result reported in Roynard, 2026). 3. Claim two — continuity belongs to the judgment layer, not the model Today an agent’s continuity lives inside the agent: in its context window, its provider-side memory features, its habits. So when context compacts, the model gets swapped, or the provider changes an API, the self tears. I’ve watched it happen mid-task, and it’s the problem that started this whole project. I flip it. The model is stateless by design — amnesiac on purpose. All provider-side memory features are off, or routed through the judgment layer. Every turn, the judgment layer assembles the working context — the stable wisdom block first (which happens to be cache-optimal), then the focus frame, then need-driven recall with provenance tags — hands it to whatever model is in the seat, and judges what of the output gets written back. The model borrows a self. It never owns one. The payoff is that the catastrophic events become non-events. Context compaction mid-thought: harmless, because the raw record and the focus tree live outside the window. Model deprecation: a config change — the next model boots into the same assembled self and can’t tell it’s a different instance. The way I put it in the design sessions: the model is this week’s cells; the judgment layer is the self. There’s a security win buried in there too. Recalled memory is injected as data with provenance, never as instructions. The only channel where the past gets to direct behaviour is wisdom that passed the governance gates in Section 5 — which is a structural defence against memory-poisoning, not a bolted-on filter. 4. Claim three — typed persistence needs an ignorance ledger I adopt Roynard’s four persistence semantics wholesale — supersession for knowledge (append-only, provenance-linked, never deleted), decay-unless-reinforced for observations, evidence-gated revision for wisdom, ephemerality for inference — along with his design litmus that decay is a storage-level property of observations only, while recency is a query-time heuristic. Facts don’t get less true with age. They get superseded, or they go stale and get revalidated, with bi-temporal timestamps keeping “when we learned it” separate from “when it was true.” I add two mechanisms. First, an epistemic-status ladder inside stored claims — observation, interpretation, hypothesis, conclusion — because an event and my judgment of the event are two different things, and a memory system that stores them in the same field will eventually believe its own moods. “A competitor launched X” and “this will destroy us” are not the same kind of object. Second — and as far as I can tell, new as a first-class memory primitive — an ignorance ledger. Classification has always been binary-plus-bin: accept into some store, or discard. The Stoics knew a third verdict: epochē, suspension of assent, for impressions you can’t accept or refuse yet. I make suspension a routing outcome. A load-bearing assumption made without evidence, or a question the record can’t answer, files as an open question object — with provenance, with what it’s blocking, and eventually with a link to the finding that answered it. At assembly time, open questions matching the current topic surface right alongside the relevant knowledge. What the system believes travels with what it knows it’s missing. Then the old knowns-and-unknowns quadrant (Johari Window, 1955; made famous by Rumsfeld, 2002; fourth quadrant per Žižek) stops being a briefing slide and becomes a schema. Known knowns: the knowledge and wisdom stores. Unknown knowns: what consolidation and promotion mine out of the raw record — patterns the system holds without knowing it holds them. Known unknowns: the ignorance ledger. Unknown unknowns: can’t be stored, obviously — but the brush against one is detectable. An episode landing far
Mike Norton (2026) studied this question.