Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 5, 2026European Journal Of Dental EducationOpen Access

Evaluating the Accuracy of Large Language Models in Dentistry: A Multi‐Model Study Using Clinical Questions From Turkey's Dental Specialty Exams

View Full Paper
Ask AI
Bookmark
Share

Authors

NKNihan KayaFŞFatih ŞengülPÇPeriş Çelikel

Discussion

Loading...

Member takes

Overview

Comparative evaluation reveals variable accuracy among five large language models answering dental specialty exam questions, indicating promise as educational aids.

Key Points

  • To evaluate and compare the accuracy of five large language models in answering clinical questions from Turkey's dental specialty examination (DUS).
  • Evaluated 1,026 clinical science questions from 13 Turkish dental specialty examinations (2012–2021) across five LLMs: ChatGPT-5, Gemini 3.5 Flash, Microsoft Copilot, DeepSeek-R1, and Grok 4.
  • Categorized questions by dental specialty and question type (information-based vs. case-based), scoring single-prompt outputs against official answer keys using chi-square and Z tests with Bonferroni correction.
  • Overall accuracy ranged from 79.2% to 90.6%, with Gemini 3.5 Flash (90.6%) and ChatGPT-5 (88.2%) achieving the highest performance, while Microsoft Copilot, DeepSeek-R1, and Grok 4 were significantly less accurate (p < 0.001).
  • Model accuracy differed significantly across both information-based (p < 0.001) and case-based questions (p = 0.029).
  • Accuracy peaked in Oral and Maxillofacial Surgery, Paediatric Dentistry, Periodontology, and Restorative Dentistry, but was comparatively lower in Prosthetic Dentistry, Orthodontics, and Endodontics.

Cite This Study

Kaya et al. (2026) studied this question.

synapsesocial.com/papers/6a9bd4216b95aff0620eb895https://doi.org/10.1111/eje.70295
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of ChatGPT on the Korean National Examination for Dental Hygienists2024 · 5 citations
  2. 2Desiderata for delivering NLP to accelerate healthcare AI advancement and a Mayo Clinic NLP-as-a-service implementation2019 · 147 citations
  3. 3Performance of ChatGPT 3.5 and 4 on U.S. dental examinations: the INBDE, ADAT, and DAT2024 · 43 citations
  4. 4The Use and Performance of Artificial Intelligence in Prosthodontics: A Systematic Review2021 · 140 citations
  5. 5Assessing the Performance of ChatGPT on Dentistry Specialization Exam Questions: A Comparative Study with DUS Examinees2025 · 7 citations