Controlled replication demonstrates attenuated behavioral effects of self-model loops in language-model agents, indicating modest path-dependent memory driven primarily by deliberative monologue.
An independent, clean-room replication of a self-model loop architecture for language-model agents (the "AC1 loop" described in Lark Laflamme's 2026 AC1-LLM / Laflamme-3T essays). The architecture wraps an LLM with a Bayesian belief over the agent's own interaction stance, re-injected into generation each turn alongside a hidden deliberative monologue. The originating write-ups report strong effects but omit the controls needed to separate the contribution of the state's content from the mere presence of an instruction, an honest task baseline, or the intrinsic inertia of the estimator. This work supplies those controls across six experiments: a structured-placebo ablation (E1), a clamped factorial over state x gate x monologue x strategy (E2), hidden-stance inference against an honest single-shot LLM baseline (E3), and closed-loop hysteresis driven against an OpenAI-compatible chat endpoint, with decay-free and state-decoupled null controls (E4–E6). Every tested claim reproduces in a "true-but-softer" form once the missing controls are added: the self-state is causally load-bearing chiefly through the hidden monologue and only when the posterior is confident; the loop is calibrated; hidden-stance inference buys persistence rather than accuracy; and the closed loop shows genuine but modest path-dependent memory (~0.07 loop area beyond accumulator arithmetic) under a single attractor, with no bistability. Scope is strictly the loop's measurable behaviour; no claim is made about consciousness or any Psi-threshold. Code and raw result data are released (MIT).
No takes yet. Share an insight, caveat, or question.
Grant Wilson (2026) studied this question.