This study looks at AI memory using the Mem0 open source system. It asks how much stored memories actually matter in an ongoing LLM conversation. Ten conditions were tested, each replicated five times. The conditions varied in how an identical behavioral constraint was delivered. In every memory condition the rule was verifiably retrieved into context, but that retrieval did not result in perfect compliance. Stored and injected deliveries bound at 19% to 84% of the same constraint typed live. This was judged on a normalized length-binding index. The largest single loss, about 42 index points, came from the memory extractor rewriting the imperative rule as a third person description. A cautionary retrieval wrapper cost about 11 index points. Identical bytes quoted back as a prior statement bound at the level of a labeled rules block, even inside the live user turn. Provenance marking, not slot or register, accounted for the remaining premium. Within the frame tested, a truthfully marked memory did not match a live instruction. That result constrains honest memory design and shows an attack surface for memory poisoning. Findings are from one model (Gemma4:31b, locally run). All logs, scoring data, and rubrics are public.
Paul Vasholz (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: