Benchmark evaluation demonstrates variation in LLM accuracy across question formats on Turkish dental specialty exams, highlighting the need for verified educational use.
Large language models (LLMs) are increasingly used as study aids in dental education, yet their performance may vary by model, discipline, and question format. This study benchmarked LLM-based systems using Turkish Dentistry Specialization Examination (DUS) Clinical Sciences questions. A public dataset of DUS Clinical Sciences multiple-choice questions from 2012 to 2021 ( n = 1,027; excluding officially canceled items) was compiled from the national examination authority. Each model was evaluated once per question: ChatGPT-4o, ChatGPT-o4-mini-high, ChatGPT-5 Thinking, Gemini-2.5 Pro, Gemini-2.0 Flash, Claude Sonnet 4, Copilot, and DeepSeek-v3.1. Accuracy was scored against the official answer key. Questions were categorized as standard, K-type, or visual. Paired question-level analyses and 95% confidence intervals were used for accuracy outcomes; response time was recorded on the same endodontics question set. Overall accuracy exceeded 80% for all models, and paired analysis confirmed significant differences among models (Cochran’s Q(7) = 247.27; p < 0.001). ChatGPT-5 Thinking achieved the highest accuracy (94.26%), followed by Gemini-2.5 Pro (92.50%); DeepSeek-v3.1 had the lowest accuracy (80.92%). Accuracy differed markedly by question format (chi-square(2) = 391.690; p < 0.001): 89.61% for standard, 75.71% for K-type, and 55.56% for visual items. Response times differed across models (Kruskal-Wallis H(7) = 649.243; p < 0.001): Claude Sonnet 4 and Copilot were fastest, while Gemini-2.5 Pro was slowest. Publicly available LLM-based systems can perform strongly on DUS-level questions, but performance is sensitive to model choice, discipline, question format, and time cost. Outputs, especially for visual and K-type items, should be used as supportive educational material and verified with reliable sources.
No takes yet. Share an insight, caveat, or question.
Şahin et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: