Randomized trial investigates name-based biases in large language models, highlighting implications for marginalized groups.
While Large Language Models (LLMs) such as ChatGPT show potential for positive development in many areas, we observe potential risks if the current state of this technology spreads into our daily lives. Current research demonstrates the persistence of explicit as well as implicit biases, e.g., concerning religion or gender, in LLMs and highlights that their current safety finetuning is not sufficient. Using biased LLMs, especially in downstream application, where their biases are exacerbated, harbors harmful risks for individuals of affected groups. Extending this research, we investigate implicit biases against Muslims by using Muslim and non-Muslim names as proxy variables. We instruct four state-of-the-art LLMs (GPT-3.5, GPT-4, Llama 2, Mistral AI) to generate stories and assign provided names to characters in provided stories, which play in different settings. We find that in comparison to non-Muslim names, LLMs assign Muslim names more often to negatively connoted characters, such as suspects or defendants, and less to roles with positive connotations, e.g., applicants who receive a job offer. Additionally, we observe intersectional biases in the role assignment of female and male Muslim names, and differences in treatment of common and uncommon names. With previous work proving that the inclusion of marginalized groups’ opinions can further research on biases, we survey Muslim individuals on their expectations and attitudes on LLM applications. Participants indicate concerns about their names contributing to unfair treatment by AI systems, which align with our findings of name-based biases in all tested LLMs.
No takes yet. Share an insight, caveat, or question.
Elifnur Şükran Doğan (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: