Experimental assay uncovers prospective self-prediction and selective revision in an LLM agent system, demonstrating grounded self-model formation.
Can an LLM-based agent system acquire a branch-specific model of its own future reactions from controlled experience, use that model prospectively, and revise it after counterevidence? We introduce a counterfactual-twin assay in which matched agent branches share the same Actor, initial conditions, available actions, and outcome multiset but differ in the contextual binding of those outcomes. The assay separates five questions that are often conflated: experience-conditioned reaction divergence, storage in a prospective Core, transfer beyond cue identity, Actor access, and selective revision. Across five precommitted source–twin pairs, progressively stronger shortcut-removal assays preserved low error for the correct Core while twin-swapped and no-Core controls remained substantially worse. In the strongest Actor assay, a Qwen3-8B Actor directly predicted ordinal changes on four reaction axes with 94.38% directional accuracy under the correct Core, compared with 0% after twin swapping, 51.88% after targeted local inversion, and 26.88% with no Core; predictions were committed before independent held-out Native Reactions were generated. A separate zero-Actor-call track tested system-level online formation and revision. Online Core formation reduced held-out binary error from 0.4875 with an empty Core to 0 with the correct learned Core, whereas the twin Core yielded error 1.0. Local counterevidence revision reduced anchor-near error from 0.41484 to 0 without changing far probes; shuffled and twin updates failed. A first-attempt confirmatory run on five precommitted unused pairs reproduced the full functional J–M–E1–F1 chain, with all five revision gains positive (mean 0.41189; exact one-sided sign-test p = .03125), 99.3% of the original mean gain. These results support a bounded claim: the system implements an experience-grounded, branch-specific prospective self-model that can guide ordinal self-prediction and undergo engineered local revision. They do not establish consciousness, phenomenological selfhood, privileged introspection, or autonomous discovery of the reaction-generating mechanism.
No takes yet. Share an insight, caveat, or question.
Shingo Akeno (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: