PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 17, 2026SWISS DENTAL JOURNAL SSO – Science and Clinical Topics0 citationsOpen Access

Leading large language models on a periodontology knowledge test

View Full Paper
ARAna-Maria RusaPSP. R. SchmidlinPSParthib Sarkar

Key Points

  • The aim is to evaluate the performance of various large language models in answering periodontology knowledge questions.
  • Evaluated four large language models: two general-purpose and two research-focused.
  • Used a validated set of 50 multiple-choice questions to assess knowledge in periodontology.
  • Conducted five independent trials for each model under both primed and non-primed conditions.
  • Analyzed performance using one-way and two-way ANOVA and independent-samples t-tests.
  • Overall accuracy across all models was 65% with a standard deviation of 3.0.
  • No significant differences in performance between models (p = 0.336).
  • Role-specific priming did not enhance model performance (p = 0.836).
  • Certain questions were consistently answered incorrectly, indicating gaps in detailed knowledge.

Abstract

Large language models (LLMs) are increasingly used in clinical and educational settings. However, there is a paucity of data on LLMs’ performance in specialized dental domains. This study assessed the performance of four LLMs, including two general-purpose models, ChatGPT-4o and DeepSeek-R1, and two research-focused models, Consensus and Perplexity, using a validated set of 50 multiple-choice questions in periodontology. Each LLM completed five independent trials encompassing the full question set under both primed and non-primed conditions. A validated answer key served as the benchmark. Performance was analyzed using one-way and two-way analysis of variance and independent-samples t-tests, with additional item-level analyses to identify questions that were consistently difficult. Overall accuracy across all models and trials was 65.0% (95% confidence interval: 63.4-66.6%) with a standard deviation of 3.0. There were no significant differences between models (p = 0.336). Role-specific priming, in which models were instructed to respond as board-certified periodontists, did not improve performance (p = 0.836). At the item level, four questions were never answered correctly, and several others were answered correctly in fewer than 13% of trials. These difficult items generally required detailed procedural knowledge, rare factual recall, or application of classification frameworks. Overall, these findings suggest that current LLMs demonstrate moderate domain knowledge in periodontology but fall short of the reliability required for unsupervised clinical decision support.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rusa et al. (2026) studied this question.

synapsesocial.com/papers/696b2696d2a12237a9349cebhttps://doi.org/10.61872/sdj-2025-04-04
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of Large Language Models on Official Periodontology Questions: A 13-Year Analysis of the Turkish Dental Specialization Examination2026
  2. 2Performance of large language models in a high-stakes dental assessment: evidence from the Turkish dentistry specialization examination2026
  3. 3Multidimensional evaluation of large language models on the AAP in-service examination: Assessing accuracy, calibration, and citation reliability2026
  4. 4A comparative analysis of the performance of leading large language models on the endodontics section of the dentistry specialization exam in Türkiye2026
  5. 5Evaluating large language models using national endodontic specialty examination questions: are they ready for real-world dentistry?2025 · 20 citations