PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 19, 2025Journal of Advances in Medicine and Medical Research0 citations

Comparative Evaluation of Multiplatform AI Performance on Practical Ophthalmology Exam Questions: Insights from the Brazilian Council of Ophthalmology Exam

View Full Paper
DNDéborah Silva NunesJDJoacy Pedro Franco DavidJFJosé Jesu Sisnando D’Araújo Filho

Key Points

  • The Gemini model achieved the highest accuracy rate of 77.6% on the ophthalmology exam questions, indicating its effectiveness.
  • Cohen's Kappa coefficient measured the agreement between AI responses and the official key, with significant performance variation observed among thematic blocks.
  • The study analyzed performances of 5 AI models on 560 exam questions from 8 ophthalmology themes, underscoring the need for supervision in AI use.
  • Findings suggest that AI can effectively assist in medical training, but caution is warranted in subjective areas requiring clinical judgment.

Abstract

In recent years, advances in artificial intelligence (AI), especially with the emergence of natural language models and deep neural networks, have revolutionised medical practice, offering tools with the potential to assist both in diagnosis and specialised medical training. The main objective of this study was to evaluate the accuracy and agreement of different artificial intelligence (AI) models in solving practical questions from the Brazilian Council of Ophthalmology (CBO) Exam. To this end, the performances of 5 AI models (ChatGPT, Gemini, DeepSeek, Google AI Studio, and GROK) were analyzed in a set of 560 questions, distributed in eight thematic blocks of ophthalmology (Cornea, Cataract, Retina, Glaucoma, Neuro-ophthalmology, Optics and Refraction, Strabismus, and Plastic Surgery/Lacrimal Duct/Orbit). The answers were compared to the official answer key by calculating the percentage of correct answers and the Cohen's Kappa and Fleiss's coefficients of agreement. Cohen's Kappa coefficient was used to measure the agreement between the AI responses and the official template, as well as Fleiss's Kappa to measure the overall agreement between the different AIs. The most evident finding was that the Gemini model presented the highest accuracy rate (77.6%) and the highest overall agreement with the official answer key. Significant variation in performance between blocks was also observed, with greater accuracy in the Retina and Glaucoma themes, and lower accuracy in the Strabismus and Plastic Surgery blocks. The thematic analysis allowed us to identify the pattern of correct answers by speciality, revealing weaknesses of the models in areas with greater dependence on visual assessment and clinical subjectivity. In addition to a probable educational applicability of AIs, it proved to be viable as a complementary tool in medical training, especially when used under supervision and with defined pedagogical objectives. Therefore, it was concluded that, despite the limitations, the most up-to-date models trained based on specific clinical data were able to faithfully reproduce diagnostic reasoning in several areas of ophthalmology, evidencing their potential for integration into specialised education, as long as they are used with technical and ethical criteria. These findings suggest AI can serve as a supplementary tool in ophthalmic education, with caution in subjective specialities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nunes et al. (2025) studied this question.

synapsesocial.com/papers/68af4953ad7bf08b1ead4fc8https://doi.org/10.9734/jammr/2025/v37i85913
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparative Assessment of Large Language Models in Optics and Refractive Surgery: Performance on Multiple-Choice Questions2025 · 1 citations
  2. 2The Role of Artificial Intelligence in Teaching Ophthalmology Skills: A Systematic Review2026
  3. 3The Performance of Artificial Intelligence-based Large Language Models on Ophthalmology-related Questions in Swedish Proficiency Test for Medicine: ChatGPT-4 omni vs Gemini 1.5 Pro2024 · 20 citations
  4. 4Artificial Versus Human Intelligence in the Diagnostic Approach of Ophthalmic Case Scenarios: A Qualitative Evaluation of Performance and Consistency2024 · 3 citations
  5. 5Comparative assessment of the accuracy of different artificial intelligence models in answering analytical and knowledge-based questions in oral and maxillofacial radiology and oral and maxillofacial surgery; a research article2026