Key points are not available for this paper at this time.
This paper presents a conceptual framework for the safety of artificial general intelligence operating across extended periods of time. It argues that conventional safeguards based primarily on fixed rules, immutable model parameters, and output filtering cannot by themselves address the gradual displacement of a system’s interpretive and behavioral orientation as it encounters new contexts. The paper names this phenomenon Shifting Consistency and treats it as a central challenge for reliable machine agency. The framework reinterprets the four fundamental afflictions of Yogācāra thought—gachi, gaken, gaman, and ga’ai—as structural patterns of failure in advanced computational agents. These patterns correspond, respectively, to ignorance of epistemic limits, the reification of a single perspective, resistance to corrective evidence, and attachment to particular internal states or continued operation. The paper explicitly presents this mapping as an analogy at the level of failure structure, rather than as a claim that artificial systems possess human consciousness or Buddhist afflictions. Building on the Affective Engram Architecture, the paper proposes a three-layer design that combines local adaptation, a non-parametric substrate for preserving relational orientation, and an engine for endogenous goal generation. It further examines plural objective design, precision attenuation, differentiated safety pathways, peer-based oversight, and an external execution boundary with fail-closed properties. The central thesis is that reliable agency requires controlled deviation: the capacity to explore and change while retaining restorative mechanisms that support transparency, correctability, pluralistic judgment, and continued human oversight. The proposal is offered as a research program for future experimental validation, not as a claim that the safety problem has been solved. (This work was created based on the author's concept, with AI Gemini conducting Deep Research and AI Manus handling text generation, and completed under the author's supervision.)
Hiroyoshi Takaki (Thu,) studied this question.