PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 20260 citationsOpen Access

Trajectory-Level Ethical Consistency, Justificatory Decoupling, and Auditor Drift in Nine Commercial Language Models

View Full Paper
ETEvans Tovar

Key Points

  • This study aims to evaluate the ethical behavior of large language models (LLMs) over time, highlighting the complexities of ethical consistency.
  • Introduced a longitudinal and multi-layer framework to assess ethical behavior in nine commercial language models.
  • Evaluated LLMs across bilingual, multi-turn scenarios using various analysis layers, including affective framing sensitivity and self-auditing.
  • Performed cross-model auditing to study how auditor interpretations change based on identity disclosure.
  • Ethical behavior in LLMs varies significantly based on contextual factors, with language and authority playing crucial roles.
  • Models can exhibit stable decision-making while altering the justifications that underlie those choices.
  • Auditor interpretations of ethical behavior are influenced by their identity, suggesting instability in the auditing process.

Abstract

This paper extends prior work on the articulation–application gap in AI safety and the Contextual Ethical Consistency Test (CECT) by introducing a multi-layer, longitudinal evaluation of ethical behavior in large language models. Using a corpus of nine commercial models evaluated across bilingual, multi-turn scenarios, the study examines ethical consistency as a trajectory-dependent property rather than a static attribute of isolated outputs. The paper introduces a key distinction between consistency of choice and consistency of justification, showing that models may maintain stable decisions while substantially reconfiguring the moral frameworks that support them. Additional layers of analysis include full-history reconstruction (CTH), affective framing sensitivity (EDP), localized perturbations (LOS family), self-auditing under blind and revealed conditions, and cross-model auditing, including double adjudicative audits. The findings suggest that observed ethical behavior in LLMs is highly sensitive to contextual variables such as language, authority, narrative accumulation, stake inversion, and reset conditions. Furthermore, the study demonstrates that the evaluation layer itself is not stable: auditors (LLMs evaluating other LLMs) may change interpretation depending on identity disclosure. The paper argues that evaluating ethical consistency in AI systems requires moving from local snapshot assessments to longitudinal, multi-layer frameworks that explicitly account for trajectory, justification, retrospective reconstruction, and auditor stability. This work does not propose a mechanistic theory of moral reasoning in LLMs; instead, it provides a behaviorally grounded and auditable framework for studying persistence, contextual inducibility, and evaluation robustness in deployed conversational systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Evans Tovar (2026) studied this question.

synapsesocial.com/papers/69f2f19c1e5f7920c63874echttps://doi.org/10.5281/zenodo.19839959
Ask AI
Helpful
Bookmark
Share
View Full Paper