PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025BMC Medical Education0 citationsOpen Access

The performance of ChatGPT on medical image-based assessments and implications for medical education

View Full Paper
YXYang XiangWCWei Chen

Key Points

  • GPT-4o achieved an accuracy of 89.5% on medical image questions, while GPT-4 scored 73.4%.
  • Prompt engineering was applied to enhance responses from both GPT-4 and GPT-4o in image assessments.
  • The study explored the application of these generative AI models in case-based teaching scenarios in medical education.
  • Human oversight is essential to ensure accuracy in AI-generated content, despite the promising performance of the models.

Abstract

Generative artificial intelligence (AI) tools like ChatGPT (OpenAI) have garnered significant attention for their potential in fields such as medical education; however, their performance of large language and vision models on medical test items involving images remains underexplored, limiting their broader educational utility. This study aims to evaluate the performance of GPT-4 and GPT-4 Omni (GPT-4o), accessed via the ChatGPT platform, on image-based United States Medical Licensing Examination (USMLE) sample items, to explore their implications for medical education. We identified all image-based questions from the USMLE Step 1 and Step 2 Clinical Knowledge sample item sets. Prompt engineering techniques were applied to generate responses from GPT-4 and GPT-4o. Each model was independently tested, with accuracy calculated based on the proportion of correct answers. In addition, we explored the application of these models in case-based teaching scenarios involving medical images. A total of 38 image-based questions spanning multiple medical disciplines-including dermatology, cardiology, and gastroenterology-were included in the analysis. GPT-4 achieved an accuracy rate of 73.4% (95% CI, 57.0% to 85.5%), while GPT-4o outperformed it with an accuracy of 89.5% (95% CI, 74.4% to 96.1%), with a numerically higher accuracy but no statistically significant difference (P = 0.137). The two models showed substantial disagreement in their classification of question complexity. In exploratory case-based teaching scenarios, GPT-4o was able to analyze and revise incorrect responses with logical reasoning. Moreover, it demonstrated potential to assist educators in designing structured lesson plans focused on core clinical knowledge areas, though human oversight remained essential. This study demonstrates that GPT models can accurately answer image-based medical examination questions, with GPT-4o exhibiting numerically higher performance. Prompt engineering further enables their use in instructional planning. While these models hold promise for enhancing medical education, expert supervision remains critical to ensure the accuracy and reliability of AI-generated content.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xiang et al. (2025) studied this question.

synapsesocial.com/papers/68c1d01a54b1d3bfb60f635ehttps://doi.org/10.1186/s12909-025-07752-0
Ask AI
Helpful
Bookmark
Share
View Full Paper