PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 18, 2025International Journal of Paediatric Dentistry0 citationsOpen Access

Challenges and Limitations of Multimodal Large Language Models in Interpreting Pediatric Panoramic Radiographs

View Full Paper
YMYuichi MineYIYuko IwamotoSOShota Okazaki

Key Points

  • Both models demonstrated significant challenges in reliable detection of teeth in pediatric dental images.
  • Claude 3.5 Sonnet exhibited higher sensitivity but lower specificity at 29.8% ± 21.5%, leading to numerous false positives.
  • OpenAI o1 outperformed Claude 3.5 Sonnet in specificity but failed to identify subtle defects in mixed dentition.
  • Variability in model results indicates the need for further refinement before clinical implementation.

Abstract

ABSTRACT Background Multimodal large language models (LLMs) have potential for medical image analysis, yet their reliability for pediatric panoramic radiographs remains uncertain. Aim This study evaluated two multimodal LLMs (OpenAI o1, Claude 3.5 Sonnet) for detecting and counting teeth (including tooth germs) on pediatric panoramic radiographs. Design Eighty‐seven pediatric panoramic radiographs from an open‐source data set were analyzed. Two pediatric dentists annotated the presence or absence of each potential tooth position. Each image was processed five times by the LLMs using identical prompts, and the results were compared with the expert annotations. Standard performance metrics and Fleiss' kappa were calculated. Results Detailed examination revealed that subtle developmental stages and minor tooth loss were consistently misidentified. Claude 3.5 Sonnet had higher sensitivity but significantly lower specificity (29.8% ± 21.5%), resulting in many false positives. OpenAI o1 demonstrated superior specificity compared to Claude 3.5 Sonnet, but still failed to correctly detect subtle defects in certain mixed dentition cases. Both models showed large variability in repeated runs. Conclusion Both LLMs failed to achieve clinically acceptable performance and cannot reliably identify nuanced discrepancies critical for pediatric dentistry. Further refinements and consistency improvements are essential before routine clinical use.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mine et al. (2025) studied this question.

synapsesocial.com/papers/68d462d231b076d99fa626c3https://doi.org/10.1111/ipd.70029
Ask AI
Helpful
Bookmark
Share
View Full Paper