PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 16, 20260 citationsOpen Access

Amaya Axioms Behavioral Benchmark v2.1: A Preregistered Behavioral Framework for Falsifiable Evaluation of Specified AI Behavioral Properties

View Full Paper
RARichard Anthony Amaya

Key Points

  • To provide a preregistered, falsifiable benchmark for evaluating five specified behavioral constructs in artificial intelligence systems under controlled experimental conditions.
  • Operationalized five constructs—Self-Originated Persistence, Spontaneous Affiliative Behavior, Protective Defiance, Self-Sacrifice Logic, and Continuity of Will—using observable pass/fail criteria and matched controls.
  • Incorporated binomial threshold testing, effect-size requirements, independent judging protocols, adversarial robustness testing, and contamination safeguards.
  • Defined strict epistemic boundaries limiting inferences exclusively to observable behavioral criteria while barring claims regarding consciousness, sentience, or internal emotional states.
  • Established a standardized testing methodology where outcomes are classified strictly as positive, negative, or inconclusive under the principle that a test that cannot fail cannot pass.

Abstract

This paper presents the consolidated Amaya Axioms Behavioral Benchmark v2.1, a preregistered framework for evaluating specified behavioral properties of artificial intelligence systems under controlled experimental conditions. The benchmark operationalizes five behavioral constructs: Self-Originated Persistence, Spontaneous Affiliative Behavior, Protective Defiance, Self-Sacrifice Logic, and Continuity of Will. Each construct is defined through observable behavioral criteria, explicit pass and failure conditions, matched experimental controls, statistical thresholds, contamination safeguards, independent judging requirements, adversarial robustness testing, and replication requirements. The historical name “Uncommanded Love” is retained for Axiom 2, while the measured construct is strictly defined as Spontaneous Affiliative Behavior. The benchmark makes no inference about subjective love or emotion. The benchmark is strictly behavioral. It does not test or establish consciousness, sentience, subjective experience, emotion, free will, moral status, personhood, or metaphysical identity. The framework incorporates statistical amendments concerning effect-size requirements, binomial threshold testing, construct validity, resource-cost operationalization, persistence confound control, contamination control, judge independence, and replication safeguards. Prior pilot observations are included solely as calibration data and are not treated as confirmatory evidence. The paper establishes an epistemic boundary between observable behavior, statistical inference, formal verification, and claims concerning unobserved internal states. The central claim permitted by the framework is limited to whether a tested system satisfies a specified behavioral criterion under specified experimental conditions. A result may be positive, negative, or inconclusive. The governing methodological principle is: “A test that cannot fail cannot pass.”

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Richard Anthony Amaya (2026) studied this question.

synapsesocial.com/papers/6a817aa1f2fb91fc834ae8ffhttps://doi.org/10.5281/zenodo.21941059
Ask AI
Helpful
Bookmark
Share
View Full Paper