Artificial intelligence (AI) and large language models (LLMs) are increasingly being explored for clinical decision-making and evidence-based information retrieval in dental implantology. However, concerns remain regarding the accuracy, reliability, and clinical applicability of different AI approaches, including general-purpose, fine-tuned, and retrieval-augmented generation (RAG) models. This study aimed to compare their ability to provide accurate, evidence-based, and clinically relevant responses to implantology-related questions. Sixteen master’s-level clinical questions were presented to all four LLMs under controlled conditions. Their anonymized responses were independently evaluated by three dental implantology specialists based on three core criteria: specificity, accuracy, and clinical utility, using a 0–10 scale. Descriptive statistics were reported as medians and interquartile ranges. Differences among models were assessed primarily using Kruskal–Wallis tests, and inter-rater reliability was evaluated using Fleiss’ kappa. LLM4 consistently achieved the highest median scores across all criteria (Specificity: median 9.0, IQR 9.0–10.0; Accuracy: median 9.5, IQR 9.0–10.0; Clinical Utility: median 10.0, IQR 9.0–10.0), significantly outperforming the other models (all p < 0.001). Cohen’s d revealed large to very large effect sizes in comparisons involving LLM4, particularly against LLM2 and LLM3. LLM2 scored lowest and showed the highest variability. Inter-rater reliability was slight across all metrics (κ = 0.078–0.180), indicating moderate evaluator disagreement. Among the models tested, LLaMA 3.2 + RAG demonstrated superior and more consistent clinical performance in dental implantology. While fine-tuning provided moderate benefits, real-time access to structured guidelines significantly enhanced response quality. Despite low inter-rater agreement, these findings highlight the promise of RAG models in supporting high-complexity, evidence-based clinical domains, while emphasizing the need for standardization in AI evaluation protocols.
Alevizakos et al. (Thu,) studied this question.