Abstract Aim: Generative artificial intelligence (AI) is emerging as a tool for clinical decision support. However, its reliability in pediatric dental diagnostics and treatment planning remains underexplored. The study was aim to evaluate and compare the diagnostic and therapeutic relevance of four AI models— ChatGPT, Gemini, Grok, and Claude—on real-world pediatric dental cases. Materials and Methods: Twelve routine outpatient cases were selected, covering a range of pediatric dental treatments. Each case included intraoral images, radiographs, and a short clinical summary. Two prompts were used for each model: Prompt 1: “According to the provided information, what is the single most correct diagnosis for tooth no. _? ” Prompt 2: “According to the provided diagnosis, what is the single most correct treatment plan for tooth no. _? ” Responses were independently scored by two pediatric dentists on a relevance scale: 0 (not relevant), 0. 5 (partially relevant), and 1 (highly relevant). Results: ChatGPT achieved the highest scores for diagnosis (91. 7%) and treatment planning (83. 3%). Gemini, Claude, and Grok showed moderate performance. Kruskal–Wallis analysis found no significant differences among the models for either diagnosis (P = 0. 48) or treatment (P = 0. 83). Agreement between evaluators was substantial, with Cohen’s Kappa ranging from 0. 62–0. 68 for diagnosis and 0. 47–0. 71 for treatment planning. Conclusion: Generative AI models demonstrated comparable performance in pediatric dental decision-making, producing clinically relevant outputs under structured prompting. Although ChatGPT 3. 5 achieved the highest numerical agreement with expert judgment, the differences among models were not statistically significant. These findings suggest that such tools may serve as supportive aids to clinical reasoning rather than replacements for professional expertise.
Rengarajan et al. (Sat,) studied this question.