Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 20, 2025Selcuk Dental JournalOpen Access

Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam

View Full Paper
Ask AI
Bookmark
Share

Authors

ÖEÖmer Ekici

Discussion

Loading...

Member takes

Overview

Comparative analysis of four large language models' performance in dentistry exams, indicating strengths in basic sciences.

Key Points

  • Claude-3.5 Haiku had the highest overall correct answer rate of 92.85%, excelling particularly in basic sciences.
  • Statistical significance in performance differences indicates LLMs struggle more with clinical sciences than basic sciences.
  • Gemini-1.5 performed worst overall, highlighting the variability in LLM capabilities across exam categories.
  • Results suggest that AI-based LLMs excel in knowledge recall but struggle in clinical reasoning and interpretation tasks.

Cite This Study

Ömer Ekici (2025) studied this question.

synapsesocial.com/papers/68d469c831b076d99fa66892https://doi.org/10.15311/selcukdentj.1674113
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of large language artificial intelligence models on solving restorative dentistry and endodontics student assessments2024 · 53 citations
  2. 2Evaluating the efficacy of leading large language models in the Japanese national dental hygienist examination: A comparative analysis of ChatGPT, Bard, and Bing Chat2024 · 37 citations
  3. 3Knowledge, attitudes, and perceptions regarding the future of artificial intelligence in oral radiology in India: A survey2020 · 90 citations
  4. 4Performance of Generative Artificial Intelligence in Dental Licensing Examinations2024 · 123 citations
  5. 5Survey of Hallucination in Natural Language Generation2022 · 4,543 citations