Aim: The aim for this study was to evaluate the efficacy of ChatGPT-4o and Claude 3.5 Sonnet in diagnosing diseases from clinical case descriptions presented in various languages. Materials and Methods: The study used 50 clinical case descriptions provided by both LLMs in three languages – English, Arabic and Ukrainian. Diagnostic results were assessed by a panel of clinicians according to standard criteria, and statistical methods were employed to determine the significance of the difference between the obtained indicators. Results: The overall diagnostic accuracy for ChatGPT-4o was 62.00%±15.62%, with performance significantly differing by language: 72.00%±10.00% in English, 44.00%±12.00% in Arabic, and 70.00%±9.00% in Ukrainian. A similar pattern was observed in Claude 3.5 Sonnet, which had a mean accuracy of 60.67%±14.50%. Both models performed significantly worse in Arabic compared to English and Ukrainian (p0.05). Conclusions: No significant differences in diagnostic accuracy were found between the LLMs; rather, the discrepancies were primarily related to the languages used. In multilingual healthcare settings, such as those involving Arabic, various underlying factors–particularly language complexity, cultural context, and the type of clinical cases-affect their performance. Further research is essential to improve the capabilities of LLMs, ensuring equitable healthcare access for diverse linguistically representative populations.
Mahdaoui et al. (Mon,) studied this question.