Key points are not available for this paper at this time.
Aim: This study aimed to evaluate the performance of large language models (LLMs)—ChatGPT-4.o, Google Gemini Advanced 2.0, and DeepSeek—in answering Turkish-language ophthalmology questions from the Turkish Medical Specialty Examination (TUS), and to compare their accuracy rates with those of ophthalmology trainees.Methods: In this cross-sectional study, 70 multiple-choice ophthalmology questions from TUS (1987–2021) were selected and categorized into six main topics: cornea-ocular surface, lens, retina, glaucoma, neuro-ophthalmology, and general ophthalmology. A total of 19 ophthalmology research assistants were grouped by their residency year. The same questions were presented to the trainees and to the three LLMs. Accuracy rates were calculated and compared using the chi-square test.Results: Among the trainees, the highest accuracy rate was observed in the third-year group, with 97.1%, while Gemini and DeepSeek stood out with 92.9% accuracy among the LLMs, ChatGPT-4.o achieved 85.7% accuracy. In the lens diseases group, the LLMs performed the lowest, whereas in retina diseases, the trainees showed the lowest performance. All LLMs achieved an overall accuracy rate of above 85%, with no significant difference between them (p0.05).Conclusion: LLMs can achieve high accuracy rates in ophthalmology exam questions and provide results close to human performance. These findings suggest that these models could be used as helpful tools in medical education and evaluation processes. However, to ensure the safe use of these models in clinical settings, their accuracy needs to be tested in studies with larger participation, and they must be specifically developed for each specialty.
Erdağ et al. (Sat,) studied this question.