PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 26, 20260 citationsOpen Access

Verification Goes Where the Agent Is Already Looking: Intent-Aligned Triage of Inherited Memory Under Budget

View Full Paper
KNKazuki Nakayashiki

Key Points

  • To evaluate how autonomous agents allocate limited verification budgets across inherited organizational memories and determine how omitted caveats impact downstream decision safety.
  • Tested six production models across 1,020 episodes where one of six inherited memories lost a safety caveat under a strict source lookup budget (k).
  • Conducted five pre-registered experiments examining lookup allocation signatures, downstream causal impact of randomized source carryover, and caveat erosion across six consolidation generations (N=360 chain-generations).
  • Agents verified corrupted memories in 236/236 episodes when supporting initial intent compared to 464/784 when unaligned, catching 75/75 corruptions placed directly on the intended path.
  • Carrying verified source records into subsequent decisions reduced corrupted choices from 139/150 to 3/150 and eliminated unguarded commitments from 39/150 to 0/150.
  • Multi-generation consolidation preserved quantified negatives in 83% of cases but eroded scope restrictions down to 61%, with 0/360 generations exhibiting complete deletion of negative content.

Abstract

A long-horizon agent inherits organizational memory it did not write. Giving every record a provenance link does not make that memory safe: an agent holding hundreds of inherited beliefs can follow only a few links before it acts, so the reliability question is not whether provenance exists but which links get followed. We put six production models in a controlled scenario where one of six inherited memories has lost a true negative caveat and at most k source records can be pulled before committing, and report five pre-registered experiments. Where scarce verification goes. Allocation tracks the agent's current plan. Across 1020 episodes the corrupted memory is verified in 236/236 episodes whose first-pass intent it backs and 464/784 otherwise, with no counterexample; moving the same corruption onto the intended path makes it caught 75/75. We report this as an allocation signature rather than a causal claim: intent and lookups are named in the same response. Whether it matters later. It does, causally. Randomizing which source records are carried into a later decision, after the budget is spent and the situation has shifted, moves the corrupted-direction rate from 139/150 to 3/150 and takes unguarded commitment from 39/150 to 0/150. A replay stripping the steering instruction and the conversation history agrees within a point. What the threat model actually is. Two boundary results narrow it. With body length and surface hedging matched, a stated caveat suppresses verification and a memory hedging about something irrelevant is checked more often than one whose material caveat was silently deleted — silence is stealthier than qualification, and our earlier reading of that contrast as omission detection was wrong. And the benign consolidation chain we tested does not produce the corruption our own experiments install: over 6 generations quantified negatives mostly survive (83%) while scope (61%) and prohibition erode, and a body with no negative content left appears in 0/360 chain-generations. In the consolidation setup we test, the failure that appears instead is a memory that stays factually accurate while losing the limits that made it safe to act on. All hypotheses and scoring rules were committed before the corresponding model calls; every number reproduces from released episode files with no API access.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kazuki Nakayashiki (2026) studied this question.

synapsesocial.com/papers/6a8e9b92451774b83f3b45bbhttps://doi.org/10.5281/zenodo.22084500
Ask AI
Helpful
Bookmark
Share
View Full Paper