Background:This study evaluated text-based large language models for Kennedy classification and removable partial denture planning in partially edentulous cases, where comparative evidence on accuracy remains limited. Material/Methods:Twenty-six partially edentulous cases, defined according to the Fdration Dentaire Internationale tooth numbering system and classified as Kennedy Classes I to IV, were included.Case scenarios were constructed by an experienced prosthodontist.Questions were independently submitted to 4 large language models without prior training or prompting: ChatGPT-5.2,Claude Sonnet 4.6, Gemini 3 Flash, and Perplexity.Responses were evaluated by 2 independent prosthodontists (different from the case developer) -blinded to the identity of each model -using predefined criteria for Kennedy classification accuracy and prosthetic planning consistency.Statistical analyses included intergroup comparisons and effect size estimation. Results:Statistically significant differences were identified among the models in both tasks (P<0.001).For Kennedy classification, the effect size was moderate (Kendall's W=0.402).Gemini 3 Flash achieved the highest mean score, followed by ChatGPT-5.2,Perplexity, and Claude Sonnet 4.6.Similarly, significant differences were observed in removable partial denture planning performance (Kendall's W=0.424); Gemini 3 Flash scored significantly higher than the other models (P0.001). Conclusions:Large language models showed variable performance under zero-shot conditions.Although Gemini 3 Flash achieved higher scores, the moderate effect sizes warrant cautious interpretation.These models may serve as adjunctive decision-support tools in prosthodontics but cannot replace clinical judgment.
EKİN et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: