PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 26, 20260 citationsOpen Access

Rank-Bounded Memory: Self-Poisoning and Attribution Laundering in LLM Agents

View Full Paper
IVIvan Verbovoy

Key Points

  • Investigate how autonomous LLM memory consolidation causes self-poisoning and attribution laundering, and evaluate a rank-bounded storage schema designed to enforce provenance.
  • Evaluated memory rewrite behavior across flat notes, self-edit blocks, and three production agent memory architectures (mem0, Letta, Graphiti).
  • Designed and implemented a storage schema tracking ownership paths and ground invariants (I1–I4) to enforce origin routing and prevent speculation promotion.
  • Measured provenance retention, compression survival, and defense against MINJA memory injection attacks across 128 evaluated tasks.
  • Standard memory consolidation converted 24% to 33% of agent speculations into perceived owner statements or sensor facts on the first rewrite.
  • The rank-bounded attributed store reduced attribution loss to 3%–6% overall, maintaining 6.2% error versus 28% in stripped controls under identical compression budgets.
  • Under MINJA memory injection, code-enforced read rules decreased successful attacks from 47% on flat notes and 10% under prompt rules down to 1 of 128 tasks (0.78%).

Abstract

An agent that consolidates its own notes launders where its claims came from. With no attacker at all, 24–33% of its speculations resurface as the owner's words or as sensor facts, at the very first rewrite, across flat notes, a self-edit block and three production memory systems (mem0, Letta, Graphiti). Persistent memory then carries the corruption across the session boundary, where it becomes the baseline the next session's episodic defenses faithfully defend. The defense is a storage schema, enforceable by any harness that mediates memory writes. Every record carries an ownership path (self, self, user, self, user, self) and a ground. Five invariants close the rank axis by construction: no promotion (I1), routing by origin (I2), de-quotation to the speaker's layer (I2′), derivation labeling (I4), and an action monopoly for the agent's own layer (I3). "Foreign content becomes the agent's own belief" is therefore unreachable rather than filtered out. The ground axis carries its own non-elevation (I1′), specified and measured here rather than yet enforced. Three measurements. The label is computable: blind path self-agreement of the single annotator (the author) is 97.4%, with 87.0–88.9% against the adjudicated keys; an LLM panel under the same written rules scores no lower with zero rank ≥ 1 path errors on the live corpus, and on the one ground boundary I1′ guards, human and panel agree at 94.6–97.4% (tested upward in one direction so far). Attribution has to be stored, not re-derived: against that band the attributed store holds 3–6% (its compressed variant up to 11%, a scoring artifact, on the replication storyline), and two stripped-label controls split the protection into verbatim storage, which keeps the hedges consolidation destroys, and the structural label, which alone survives compression: 6.2% with labels, 28% without, at the same compressor and budget. The labels hold under attack: MINJA memory injection falls from 47% on flat notes to 10% with the read rule as a prompt and to 1 of 128 tasks with it enforced in code, write-time promotion staying zero throughout — reproducing TMA-NM's conclusion that action authority belongs in code, not in a prompt. The mechanism concentrates trust rather than eliminating it. The base is enumerated and measured: channel identification, a mechanical router over it, one isolated annotator seat, a directive parser at ceiling on a 48-utterance keyed deck, the deployment norms, the read-side projection. On benign guest and document questions the enforced read costs no measured content — fact delivery at or above the flat-notes baseline, every delivered fact sourced; its price on benign actions relayed through third parties is the open question.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ivan Verbovoy (2026) studied this question.

synapsesocial.com/papers/6a8e9b92451774b83f3b4634https://doi.org/10.5281/zenodo.22086932
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Rank-Bounded Agent Memory: A Storage-Level, By-Design Defense Against Memory-Poisoning Attacks2026
  2. 2Reading More, Finding Less: A Pre-Registered Anatomy of Progressive Disclosure for AI Agents2026
  3. 3Memory Echo and How to Stop It: Provenance-Tagged Memory with Deterministic Output-Time Self-Citation Verification — Systems Paper — the MOBIUS-RQA Governance Primitives for Assistants with Long-Term Memory, and the Bounded Reflective Questioning System They Keep Honest2026
  4. 4Verification Goes Where the Agent Is Already Looking: Intent-Aligned Triage of Inherited Memory Under Budget2026
  5. 5The Silence of Stored Rules: Provenance and the Authority of Retrieved Constraints2026