Randomized trial verifies self-created objectives in artificial agents, indicating strong detection capabilities.
This record is a three-manuscript research package, released together and intended to be read as one unit. PAPER I — Verifying Self-Created Objectives I: A Proposer-Agnostic, Pre-Registered Falsification Protocol with a Measured Adversarial False-Accept Rate (46 pages) A pre-registered, proposer-agnostic falsification protocol of six controls for self-created objectives: five test grounding (scramble, ablation, a matched decoy, unprompted restraint, attribution) and a sixth, C6, tests maintained-as-end behavior. The protocol is instantiated on EVE, a homeostatic drive system, and on a second public agent, with independent external reimplementation and reproduction of the emergence phenomenon (Studies 0b–0d). Evolutionary, gradient, and RL attackers over ~5×10⁴ candidates per space measure a C1–C5 false-accept rate of 7×10⁻⁴; with C6 added and re-frozen, the adaptive searches accept none (0/51,200), and a same-budget i.i.d. arm licenses a rule-of-three 95% upper bound of ≤5.9×10⁻⁵. Every soundness rate is conditional on the authored candidate grammar; scope and named exclusions are stated in the paper. PAPER II — Verifying Self-Created Objectives II: Endogenous Self-Extension on a Deployed Substrate: a Pre-Registered Gate, Blind External Forecasts, and a Discovery (17 pages) Deploys the protocol. Study 5: a pre-registered gate on the deployed substrate, with no per-run relaxation, under which an un-authored uncovered consequence arises across three construction paths, passes all six controls to bench standard, persists as a reused drive, and reproduces on a second agent (HomeoGrid) and in two cross-validated external environments, against 0/50 on a matched no-stake control. Study 6: blind external forecasts — two external parties cross-validate in swapped author/evaluator roles, each registering a forward behavioral prediction before unsealing, the signed digests blockchain-anchored; the verified objective forecasts novel behavior in conditions never seen, while a no-objective agent fails the identical test. Study 7: on 10 world-instances whose coupling structures an external party draws from a distribution the author never sees, certified self-created objectives are confirmed by an out-of-loop analyst to track novel true regularities of the world — of 9 certified by C1–C6, 7 are discoveries on no pre-run list, with zero false discoveries. STANDALONE LETTER — A Measured False-Accept Rate for a Proposer-Agnostic Falsification Battery (9 pages) The headline quantity in standalone form, for a methods audience: over a defined candidate space, the five grounding controls C1–C5 admit a measured false-accept rate of 7.2×10⁻⁴ (37/51,200), every slip a single grounded-but-non-end class that C6 rejects; the full C1–C6 conjunction accepts 0/51,200. A double-blind detection study measures sensitivity 0.95 (38/40) at a false-positive rate of 0.05 (1/20). The claim is a measured rate over a defined synthetic space, not that the protocol is sound, and not that it detects mis-generalized objectives inside a trained policy. HOW TO READ THE PACKAGE. Paper I carries the protocol, the control benches, and the formal characterization; Paper II carries the deployment studies; the letter issues the measured rate standalone. Every constant, threshold, and predicted verdict is SHA-256 hash-locked before the corresponding run. Funding: internal funds of Qualion Intelligence, a research programme of AWELON d.o.o. (Slovenia); no external funding. Author: Matija Ludvig, Qualion Intelligence (ORCID 0009-0000-6863-9474). Contact: contact@qualion-intelligence.com.
No takes yet. Share an insight, caveat, or question.
Matija Ludvig (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: