Key result
GPT-4.0 outperforms humans and GPT-3.5 on ophthalmology self-assessment questions with ~82% accuracy.
Why the study?
The study was conducted to compare the performance of humans, GPT-4.0, and GPT-3.5 in answering multiple-choice questions from the American Academy of Ophthalmology Basic and Clinical Science Course self-assessment program.
Cross-Sectional (n=1,023)
Absolute Event Rate: 82.4% vs 75.7%
p-value: p=<0.0001
GPT-4.0 significantly outperformed both human respondents and GPT-3.5 on an ophthalmology self-assessment test, though it struggled more with surgery-related questions.
GPT-4.0 may aid ophthalmology education; hypothesis-generating and should not yet change clinical practice or training.
To compare the performance of humans, GPT-4.0 and GPT-3.5 in answering multiple-choice questions from the American Academy of Ophthalmology (AAO) Basic and Clinical Science Course (BCSC) self-assessment program, available at https://www.aao.org/education/self-assessments. In June 2023, text-based multiple-choice questions were submitted to GPT-4.0 and GPT-3.5. The AAO provides the percentage of humans who selected the correct answer, which was analyzed for comparison. All questions were classified by 10 subspecialties and 3 practice areas (diagnostics/clinics, medical treatment, surgery). Out of 1023 questions, GPT-4.0 achieved the best score (82.4%), followed by humans (75.7%) and GPT-3.5 (65.9%), with significant difference in accuracy rates (always P < 0.0001). Both GPT-4.0 and GPT-3.5 showed the worst results in surgery-related questions (74.6% and 57.0% respectively). For difficult questions (answered incorrectly by > 50% of humans), both GPT models favorably compared to humans, without reaching significancy. The word count for answers provided by GPT-4.0 was significantly lower than those produced by GPT-3.5 (160 ± 56 and 206 ± 77 respectively, P < 0.0001); however, incorrect responses were longer (P < 0.02). GPT-4.0 represented a substantial improvement over GPT-3.5, achieving better performance than humans in an AAO BCSC self-assessment test. However, ChatGPT is still limited by inconsistency across different practice areas, especially when it comes to surgery.
No takes yet. Share an insight, caveat, or question.
Taloni et al. (2023) conducted a cross-sectional in Ophthalmology knowledge assessment (n=1,023). GPT-4.0 vs. Humans was evaluated on Percentage of correct answers on multiple-choice questions (p=<0.0001). GPT-4.0 achieved a significantly higher accuracy rate (82.4%) compared to humans (75.7%) and GPT-3.5 (65.9%) on the American Academy of Ophthalmology self-assessment multiple-choice questions.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: