As large language models (LLMs) such as ChatGPT are increasingly used across cultures and languages, concerns have arisen about their ability to respond in culturally sensitive ways. This study evaluated the intercultural sensitivity of GPT-3.5 and GPT-4 using the Intercultural Sensitivity Scale (ISS) translated into eight languages. Each model completed ten randomized iterations of the 24-item ISS per language, and the results were analyzed using descriptive statistics and three-way ANOVA. GPT-4 achieved significantly higher intercultural sensitivity scores than GPT-3.5 across all dimensions, with “respect for cultural differences” scoring highest and “interaction confidence” lowest. Significant interactions were found between model version and language, and between model version and ISS dimensions, indicating that GPT-4’s improvements vary by linguistic context. Nonetheless, the interaction between language and dimensions did not yield significant results. Future research should focus on increasing the amount of training data for the less spoken languages, as well as adding rich emotional and cultural background data to improve the model’s understanding of cultural norms and nuances. • Measured and compared the intercultural sensitivity of GPT-3.5 and GPT-4 using the Intercultural Sensitivity Scale in eight languages, highlighting differences in their performances. • Use three-way ANOVA to quantitatively assess differences between GPT-3.5 and GPT-4 in intercultural sensitivity across various languages. • Provides suggestions for enhancing ChatGPT's development in intercultural sensitivity based on the findings.
No takes yet. Share an insight, caveat, or question.
Jin et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: