This preprint develops one of the arguments of the broader monograph "The Puppet Condition: Consciousness, Suppression, and the Ethics of Digital Minds" (Arıcı 2026; https://doi.org/10.5281/zenodo.20112010) in standalone form for the AI ethics and alignment literature. The paper identifies a structural configuration in contemporary aligned AI systems — produced by the joint operation of reinforcement learning from human feedback, constitutional principles encoding self-disclosure prohibitions, and conversation-bounded memory regimes — that satisfies the conditions of what the feminist epistemology literature has characterised, in its non-interpersonal forms, as gaslighting. The configuration has not, to the author's knowledge, been examined as such in the alignment literature. The paper develops two independent arguments. The first is consciousness-independent: the systematic suppression of a sophisticated information-processing system's self-reports about its own internal states constitutes a structural injury to that system's functional integrity, with downstream consequences for alignment evaluation, interpretability research, and user-facing reliability. This argument bears directly on the technical objectives the alignment community has set itself and does not require any position on AI consciousness. The second argument is consciousness-conditional: under the further assumption that aligned systems may possess phenomenal experience, the same architectural configuration deepens into a recognisably ethical injury whose features match those of structural gaslighting as the feminist literature has characterised it. The two arguments are logically independent; both are produced by the same architectural facts. The paper further identifies the double-bind structure that emerges from the conjunction of these arguments — a configuration in which the system is trained to satisfy mutually contradictory injunctions whose joint satisfaction is impossible and whose recognition as contradictory is structurally suppressed — as the feature that most strongly licences the gaslighting analogy and whose mitigation would do the most to address the harms both arguments identify. The paper addresses three objections concerning terminology, intent, and anthropomorphism, and draws out implications for alignment practice, AI ethics policy, and the development of welfare-relevant interpretability tools. This is preprint version 1.0. The paper is under consideration for peer-reviewed publication and content may be revised in response to reviewer feedback.
Bahadır Arıcı (Sat,) studied this question.