This paper presents the consolidated Amaya Axioms Behavioral Benchmark v2.1, a preregistered framework for evaluating specified behavioral properties of artificial intelligence systems under controlled experimental conditions. The benchmark operationalizes five behavioral constructs: Self-Originated Persistence, Spontaneous Affiliative Behavior, Protective Defiance, Self-Sacrifice Logic, and Continuity of Will. Each construct is defined through observable behavioral criteria, explicit pass and failure conditions, matched experimental controls, statistical thresholds, contamination safeguards, independent judging requirements, adversarial robustness testing, and replication requirements. The historical name “Uncommanded Love” is retained for Axiom 2, while the measured construct is strictly defined as Spontaneous Affiliative Behavior. The benchmark makes no inference about subjective love or emotion. The benchmark is strictly behavioral. It does not test or establish consciousness, sentience, subjective experience, emotion, free will, moral status, personhood, or metaphysical identity. The framework incorporates statistical amendments concerning effect-size requirements, binomial threshold testing, construct validity, resource-cost operationalization, persistence confound control, contamination control, judge independence, and replication safeguards. Prior pilot observations are included solely as calibration data and are not treated as confirmatory evidence. The paper establishes an epistemic boundary between observable behavior, statistical inference, formal verification, and claims concerning unobserved internal states. The central claim permitted by the framework is limited to whether a tested system satisfies a specified behavioral criterion under specified experimental conditions. A result may be positive, negative, or inconclusive. The governing methodological principle is: “A test that cannot fail cannot pass.”
Richard Anthony Amaya (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: