Los puntos clave no están disponibles para este artículo en este momento.
Abstract Traumatic dental injuries (TDIs) are complex clinical conditions that require timely and accurate decision-making. With the rise of large language models (LLMs), there is growing interest in their potential to support dental management. This study evaluated the accuracy and consistency of DeepSeek R1's responses across all categories of TDIs and benchmarked its performance against other common LLMs. DeepSeek R1 and six other LLMs, ChatGPT-4o mini, ChatGPT-4o, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Flash, and Gemini 1.5 Advanced, were assessed using a validated question set (125 items) covering five subtopics: general introduction, fractures, luxations, avulsions of permanent teeth, and TDIs in the primary dentition (25 items per group) with a specific prompt. Each model was tested with five repetitions for all items. Accuracy was calculated as the percentage of correct responses, while consistency was measured using Fleiss' kappa analysis. Kruskal–Wallis H and Dunn's post-hoc test were applied for comparisons of three or more independent groups. DeepSeek R1 achieved the highest overall score of 86.4% ± 2.5%, despite the most inconsistent responses (κ = 0.694), statistically higher than those of ChatGPT-4o mini (74.7% ± 0.9%), Claude 3 Opus (75.2% ± 1.0%), and Gemini 1.5 Flash (73.85% ± 2.3%) (p < 0.0001). Across all models, accuracy was notably lower for luxation injury questions (68.3% ± 3.2%). LLMs achieved moderate to high accuracy, yet this was tempered by varying degrees of inconsistency, particularly in the top-performing DeepSeek model. Difficulty with complex scenarios like luxation highlights current limitations in artificial intelligence (AI)'s diagnostic reasoning. AI should be viewed as a valuable dental educational and clinical adjunctive tool for knowledge acquisition and analysis, not a replacement for clinical expertise.
Termteerapornpimol et al. (Wed,) studied this question.