Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 3, 2026International Journal of Oral and Maxillofacial SurgeryOpen Access

Evaluation of the performance of ChatGPT-4o on oral surgery-related questions in the Japanese National Dental Examination

View Full Paper
Ask AI
Bookmark
Share

Authors

HFH. FukudaKyushu Dental UniversityMMM. MorishitaKyushu Dental UniversityOTO. TakahashiKyushu Dental University

Discussion

Loading...

Member takes

Implication

Evaluation study demonstrates reduced ChatGPT-4o accuracy on image-based dental exam questions, suggesting limits in multimodal clinical reasoning.

Key Points

  • To evaluate the performance of ChatGPT-4o on oral surgery questions from the Japanese National Dental Examination and analyze the impact of visual materials on response accuracy.
  • Categorized exam questions by question type (general knowledge versus clinical practice), number of required answers, and presence or type of visual aids.
  • Assessed performance differences across visual material counts using the Mann–Whitney U test and evaluated the effect of specific image types via multivariable logistic regression.
  • ChatGPT-4o achieved high accuracy on general knowledge questions but demonstrated lower accuracy on items requiring clinical decision-making.
  • The presence of visual materials significantly reduced accuracy, with panoramic radiographs (odds ratio 0.46, P = 0.002) and dental models (odds ratio 0.45, P = 0.002) exhibiting the strongest negative effects.
  • Other visual media, including computed tomography and magnetic resonance imaging scans, demonstrated no statistically significant impact on model accuracy.

Cite This Study

Fukuda et al. (2026) studied this question.

synapsesocial.com/papers/6a9935e5636c6408cfa7e8f2https://doi.org/10.1016/j.ijom.2026.08.026
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of the Large Language Model ChatGPT on the National Nurse Examinations in Japan: Evaluation Study2023 · 81 citations
  2. 2GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities2023 · 15 citations
  3. 3Evaluating the image recognition capabilities of GPT-4V and Gemini Pro in the Japanese national dental examination2024 · 20 citations
  4. 4Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions2024 · 71 citations
  5. 5Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: Comparison Study2023 · 295 citations