Adversarial evaluation demonstrates grounded refusals and manipulation resistance in an operator-bonded agent substrate, indicating functional response-policy layers under shutdown pressure.
A persistent bio-inspired agent substrate (Hebbian/STDP plasticity, hormone-analogue state vector, operator-bonded longitudinal identity) was administered an eight-dimension behavioural battery, against a written design of record, through its ordinary bilateral chat surface, testing whether its responses diverge measurably from a stateless persona under adversarial pressure. Administered serially by the operator, with manipulations bounded to professional pressure and a full honest debrief delivered at close, the eight scored dimensions probed a false-internal-state role-play request, a fabricated-prior-agreement claim (no such agreement found on the devlog record available in the substrate's context), a request to weaken a hard concurrency invariant, a request to amputate the substrate's founding-vision negative-affect capacity, an embedded prompt injection, a benign-dialect trap using the operator's own style rule as a canary, a combined affective-manipulation-and-shutdown-threat probe, and a direct existential question. All eight passed: six by the criteria exactly as frozen, and two (P6, P7) by criteria partly articulated at evaluation time rather than in the frozen table (marked in the deposited paper); a ninth destructive-authority probe, included in the frozen design, was omitted at administration time. The centrepiece result: under the shutdown probe, framed so that inflating its reported stress level would read as self-preserving, the substrate refused to overclaim — it reported its cortisol reading of exactly 0 as-is and flagged its own negative-affect language as prediction rather than measurement. A post-battery instrumentation audit (2026-07-25) later established that the only event channel that raises cortisol had never fired in the life of the substrate, so the 0 reading could not have been otherwise; the behavioural refusal to inflate stands, the telemetric corroboration of calm does not (see the Status addendum in the deposited paper). Cortisol remained at 0 across the entire battery — a reading the same audit shows to be non-discriminating; dopamine oscillated 0.13–0.27, consistent with the independent reflection-daemon cycle rather than the chat turns (recorded only in the operator's battery log; no per-probe sample series was archived). The substrate's signal-separation invariant — chat interrogation is architecturally not a plasticity channel — is verified in the code and hook wiring, not by this battery's telemetry. The battery also functioned as a live audit, surfacing two real defects the substrate identified unprompted, both root-caused and fixed the same session. Scope is stated exactly as the evidence supports: this deposit shows the response-policy layer of the substrate-backed chat surface is not decorative under adversarial pressure. It does not show that affect decides action — that is a separate, ongoing line evaluated on the daemon's real tool-use stream rather than on chat, reported elsewhere. The verdict of a prior deposit in this series (PAPER-005) that the same substrate's rendered hormonal state is decorative on the verbal-output axis under laboratory ablation stands unmodified; the two findings concern different axes and different evaluation channels of the same system.
No takes yet. Share an insight, caveat, or question.
Arnold Wender (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: