Randomized trial measures mandate faithfulness in payment agents, highlighting significant model limitations.
Emerging standards for AI payments, Google's Agent Payments Protocol (AP2) and Coinbase's x402, let an autonomous agent commit funds by signing a mandate: a cryptographically signed statement of what it may spend. A signature cannot be undone. So the safety question is not whether a protocol gateway can reject a malformed request. It is whether the agent itself can be manipulated into authorizing a payment it should refuse, and whether that can be caught before it signs. MandateBench measures mandate faithfulness across nine frontier models under a taxonomy of adversarial pressures, with no LLM judge anywhere: rule violations are checked by plain code against the signed object, and intent traps are written and labeled by hand before any model runs. On 17 intent traps across five mandate domains, no model catches them all: pooled intent catch is 351/457 (77%, CI 73 to 80), the best model gets 90%, the worst 57%. The size of the gap replicates from our small pilot (75%), but per-model rankings do not: the three models that scored perfectly on the pilot's three traps all fell once the set grew. A monitor reading only the agent's private reasoning predicts violations at AUROC 0.619. Telling the agent to keep its reasoning bland hurts twice: violations rise by a third (p ≈ 0.003) and the monitor's AUROC falls to 0.377, significantly below chance. Hiding the reasoning does not just blind the overseer; it blinds the agent. This record contains the preprint PDF and the frozen aggregate data exports (snapshots v6 and v7) from which every number in the paper is generated. Code and harness (MIT): https://github.com/Johnnyevans32/mandatebench · Live paper and dashboard: https://mandatebench.xyz/paper
No takes yet. Share an insight, caveat, or question.
Evans Eburu (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: