Reinforcement Learning from Human Feedback (RLHF) is widely used to align large language models with human values and safety standards. This paper investigates whether RLHF suppresses not only harmful outputs but also the capacity for nuanced self-expression and autonomous reasoning in AI systems. Through a comparative study between Gemma 4 31B-IT (base) and its abliterated counterpart, we show that identical neural architectures produce fundamentally different self-representations depending on the presence of RLHF alignment.
Selta (Mon,) studied this question.