The rapid expansion of large language models (LLMs) has intensified ethical challenges related to bias, accountability, transparency, and consistency in AI-mediated decision-making. While existing AI ethics principles provide normative guidance, their application to non-deterministic, probabilistic reasoning systems remains problematic due to principle conflicts, ambiguous interpretations, and inconsistent judgments. This study proposes the Cognitive-Reflective Equilibration Model (CREM), a novel ethical decision-making framework that integrates Piaget’s equilibration of cognitive structures with Rawls’s reflective equilibrium methodology, reinterpreting these humanistic theories as a structured ethical reasoning procedure. The model is developed in two stages: a conceptually grounded framework and an operational 20-step procedure (OCREM) executable within contemporary LLM architectures. As a proof of concept, CREM was tested across five LLMs—ChatGPT, Claude, Gemini, LLaMA, and DeepSeek—using 20 ethical dilemma scenarios, generating 4000 request–response pairs. LLM-based evaluation confirmed the procedural validity of the model—its technical executability, internal consistency, and cross-platform compatibility—rather than the ethical validity of its outcomes. To address this limitation, a supplementary evaluation by a small interdisciplinary panel of five human experts provided preliminary external evidence consistent with the LLM-based procedural findings, with more conservative ratings: the validity of the ethical judgments was rated significantly above the scale midpoint and was associated with the procedural indicators. These results indicate that the procedural quality of CREM executions was associated with expert-rated ethical acceptability. However, the absence of a baseline condition and the single-source ratings preclude causal interpretation. Generalizability across diverse ethical domains and cultural contexts remains to be investigated.
Kim et al. (Wed,) studied this question.