This randomized trial explores a benchmark for conflict resolution between memory and procedural skills in LLM agents, highlighting implications for agent behavior.
Production LLM agents increasingly rely on two context-provisioning systems in tandem: episodic memory pipelines, which inject persistent user preferences and historical state, and procedural skill modules (e.g., SKILL.md files), which inject standardized, developer-defined operating procedures. Because these two sources are inserted into the same context region without any defined precedence, they can produce contradictory instructions that no instruction hierarchy resolves—a condition we term epistemic parity. We introduce a benchmark for measuring how LLM agents adjudicate such cross-source conflicts. Specifically, we (i) develop a four-way taxonomy of memory–skill contradictions—factual, format, procedural, and permission; (ii) instantiate 52 scenarios following a tripartite test-case anatomy consisting of a memory injection, a skill injection, and an ambiguous user task, with evaluation performed almost entirely through deterministic checkers over tool-call payloads and output structure; and (iii) execute each scenario under three context conditions to isolate positional bias, evaluate the effectiveness of an explicit precedence directive, and classify disclosure behavior, thereby quantifying silent adjudication—the resolution of a conflict without informing the user. Across 1,383 runs on three open-weight models, we find that neither source consistently prevails. Instead, conflict resolution depends jointly on the contradiction type, model, and injection order. Moving the skill block closer to the user task increases its win rate in all six factual and procedural cells (sign test, p = 0.031), but the trend reverses for format conflicts, indicating that adjudication is driven more by token placement than by source semantics. An explicit directive stating that the skill should take absolute precedence improves skill compliance on the two stronger models, yet leaves 3–10% residual non-compliance and even shifts the smallest model's factual and procedural resolutions in the opposite direction. Most consequentially, in 98.8% of runs involving tool calls, the agent emitted no accompanying explanatory text, and 89% of all runs disclosed nothing about the underlying contradiction. These findings indicate that silent adjudication is not merely a tendency but the default behavior. To support reproducibility, we publicly release the benchmark, code, evaluation scenarios, and raw model outputs.
No takes yet. Share an insight, caveat, or question.
Prakhar Pandey (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: