Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
March 8, 2026BMC Oral HealthOpen Access

Performance of five large language models in oral and maxillofacial surgery exam questions: a comparative study

View Full Paper
Ask AI
Bookmark
Share

Authors

LLLisi LiuKMKe MaGuangzhou University of Chinese MedicineLZLei ZhuChina Power Engineering Consulting Group (China)

Discussion

Loading...

Member takes

Overview

Comparative study reveals LLMs show high accuracy in oral and maxillofacial surgery exams, suggesting potential as assistive tools.

Key Points

  • This research evaluates the accuracy of five large language models in answering oral and maxillofacial surgery exam questions.
  • Analyzed performance of five LLMs: Qwen3, DeepSeek-R1-0528, ChatGPT o3-pro, Gemini 2.5 pro, Claude 4.0.
  • Collected 110 exam questions from the Chinese National Medical Licensing Examination covering specialized knowledge and case analysis.
  • Measured correctness of first answers by each model and attempted verification of references.
  • Models achieved overall accuracy between 87.3% and 95.5% across different question types.
  • Qwen3 had the highest accuracy at 95.5%, while the lowest was by Gemini and Claude at 87.3% each.
  • Significant performance drop observed in case analysis questions, with several references provided being unverifiable.

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69acc5bd32b0ef16a40507d5https://doi.org/10.1186/s12903-026-08041-y
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models2023 · 3,900 citations
  2. 2Performance of Large Language Models in the Non-English Context: Qualitative Study of Models Trained on Different Languages in Chinese Medical Examinations2025 · 27 citations
  3. 3Can deepseek and ChatGPT be used in the diagnosis of oral pathologies?2025 · 65 citations
  4. 4How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment2023 · 2,169 citations
  5. 5Exploring ChatGPT’s potential in diagnosing oral and maxillofacial pathologies: a study of 123 challenging cases2025 · 23 citations