Key points are not available for this paper at this time.
: Generative AI systems are increasingly deployed in decision-support settings, raising important questions about how they respond to identity-sensitive scenarios. While previous studies have primarily examined bias as a static property of AI models, comparatively little attention has been paid to how response patterns vary across different contextual conditions. This study conducts a structured audit of ChatGPT outputs by comparing multiple model versions across decision scenarios involving gender, race, religion, and disability under ambiguous and explicit contextual conditions. Using a prompt-based experimental design, we examine response distributions, neutral-response frequencies, and qualitative reasoning patterns across model versions and contexts. The results indicate that response patterns vary systematically with contextual information. Ambiguous scenarios produce substantially more neutral responses, whereas explicit scenarios more frequently lead to determinate judgments. Importantly, neutral responses are not interpreted as inherently fair or unbiased. Instead, we distinguish between appropriate neutrality, in which insufficient evidence justifies withholding judgment, and neutrality overuse, in which excessive caution leads models to avoid making contextually supported decisions. Rather than treating bias as a fixed model characteristic, this study conceptualizes AI outputs as context-dependent patterns of response expression. The proposed framework provides a structured approach for evaluating neutrality, judgment, and response variation in large language models and offers methodological guidance for future AI-output auditing research.
Hong et al. (Wed,) studied this question.