Key points are not available for this paper at this time.
Introduction: Artificial intelligence (AI) chatbots are increasingly being used in healthcare. However, their diagnostic performance in endodontics remains underexplored. This study assessed and compared the diagnostic accuracy and temporal consistency of four AI chatbots—ChatGPT 4.5, Gemini 2.5 Pro, Glass Health, and MedGebra GPT-4o—in interpreting endodontic conditions in clinical scenarios. Methods: A cross-sectional observational study was conducted using 16 peer-reviewed endodontic case reports. Each case was presented to the four chatbots with a standardized diagnostic query at three different times of day, generating 192 responses. Outputs were scored on a four-point scale against gold-standard diagnoses from the published reports. Diagnostic accuracy and consistency were assessed using Kruskal–Wallis H, Mann–Whitney U, and Friedman tests. Results: Gemini 2.5 Pro achieved the highest diagnostic accuracy (mean = 2.81) and consistency (75%), followed by Glass Health (mean = 2.35, consistency = 43.8%), ChatGPT 4.5 (mean = 1.96, consistency = 50%), and MedGebra GPT-4o (mean = 0.9, consistency = 18.8%). Significant temporal variation was observed only in Gemini 2.5 Pro ( P = 0.018), which showed reduced accuracy at noon. No significant time-of-day effect was observed for the other chatbots ( P > 0.05). Statistical tests confirmed significant inter-chatbot differences in both accuracy ( P < 0.001) and consistency ( P = 0.017). Conclusion: Gemini 2.5 Pro and ChatGPT- 4.5 demonstrated promising diagnostic accuracy for practical applications. While AI chatbot cannot replace endodontic experts, it may serve as a valuable decision-support tool for less experienced dentists, particularly in generating preliminary diagnoses based on signs and symptoms.
F Alnassar (Fri,) studied this question.