PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 2026International Endodontic Journal2 citations

Testing LLM Diagnostics in Endodontics: The Impact of Linguistic Variation on Unseen Cases

View Full Paper
IBItrat BatoolNNNighat NavedFUFahad Umer

Key Points

  • The main aim is to assess how well GPT‐5 Plus and Gemini 2.5 Flash diagnose unseen endodontic cases and how linguistic variations affect this performance.
  • Utilized a benchmark dataset of 100 clinical case MCQs.
  • Introduced linguistic variations through paraphrasing, perturbation, and permutation.
  • Evaluated model performance metrics including accuracy and F‐1 score.
  • Analyzed model agreement using Cohen's κ and conducted paired comparisons with McNemar's test.
  • GPT‐5 Plus achieved 80% accuracy while Gemini 2.5 Flash had 66% accuracy on the benchmark dataset.
  • Performance of GPT‐5 Plus declined with linguistic perturbation, showing significant negative impact.
  • Gemini 2.5 Flash demonstrated consistent decision patterns across all linguistic transformations without significant performance drops.
  • Qualitative analysis rated Gemini 2.5 Flash higher in reasoning quality for both original and varied datasets.

Abstract

ABSTRACT Aim To assess the diagnostic performance of two language models, GPT‐5 Plus and Gemini 2.5 Flash using a curated benchmark dataset of unseen endodontic and restorative dentistry related clinical case scenarios and the linguistic variations introduced around the original dataset. Additionally, a descriptive qualitative analysis was performed on a subset of cases to evaluate the quality of reasoning generated by both models. Methodology One hundred single best answer MCQs were generated using standardised resources, constituting a benchmark dataset. Controlled linguistic variations were introduced around the original dataset; paraphrasing (sentence/clause rewording), perturbation (token‐level substitutions), and permutation (answer‐order shuffle). These case scenarios were presented to both models using a standardised prompt, and the performance metrics (accuracy/recall, F‐1 score) were computed. Agreement between and within models was analysed using Cohen's κ, while paired differences were evaluated using McNemar's test with a significant p‐ value < 0.05. Qualitative analysis was performed on a subset of the total sample, and the responses were evaluated on a 3‐point Likert scale. Results GPT‐5 Plus achieved 80% accuracy on benchmark dataset compared to 66% for Gemini 2.5 Flash (McNemar's p ‐value = 0.0066). When linguistic variations were introduced, the performance of GPT‐5 Plus declined, with perturbation having the most significant effect (McNemar's p ‐value = 0.003). Gemini 2.5 Flash, on the other hand, though inferior initial performance on benchmark dataset, maintained uniform decision patterns across all transformations with no significant drop further. The descriptive qualitative analysis demonstrated an overall higher proportion of responses rated as good (8/10, 80% for original dataset; 7/10, 70% for linguistic variations) for Gemini 2.5 Flash as opposed to GPT‐5 Plus. Conclusion GPT‐5 Plus outperformed Gemini 2.5 Flash on benchmark dataset; however, it was sensitive to linguistic variations. Perturbation negatively influenced the performance of GPT‐5 Plus, emphasising the need to further investigate the linguistic phenomenon that may have affected the model's degradation. Additionally, the descriptive qualitative analysis demonstrated relatively higher performance for Gemini 2.5 Flash compared to GPT‐5 Plus on the original dataset and across linguistic variations. However, owing to the descriptive nature of findings and limited sample size, the results should be interpreted with caution.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Batool et al. (2026) studied this question.

synapsesocial.com/papers/698829520fc35cd7a884993bhttps://doi.org/10.1111/iej.70109
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparison and Review of Different Versions of <scp>OpenAI</scp> Chat <scp>GPT</scp> , Anthropic Claude and Google Gemini Large Language Models' Performance on Endodontics Questions in the Turkish Dentistry Specialization Exam2026
  2. 2Examiner stratification reveals clinically relevant variability in large language model answers to endodontic patient questions2026
  3. 3Performance of large language models on undergraduate endodontic multiple-choice questions2026
  4. 4Benchmarking GPT-5, Gemini 2.5 Pro, Grok 4, and other LLMs on pediatric dentistry questions from a dental specialization exam2026
  5. 5Effectiveness of Various General large language models in Clinical Consensus and Case Analysis in Dental Implantology: A Comparative Study2024