ABSTRACT Aim To compare retrieval‐augmented systems with general‐purpose large language models (LLMs) on standardised periodontal clinical vignettes. Materials and Methods Eleven AI systems were evaluated: nine general‐purpose LLMs, one general‐purpose retrieval‐augmented platform (Perplexity) and one medical‐domain retrieval‐augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5‐point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover–Iman tests with Holm correction were applied; mixed‐effects and ordinal models served as sensitivity analyses. Results At least one parameter scored dangerous (≤ 2) in 3.3%–46.7% of responses across platforms, despite mean composite scores (3.28–4.86) exceeding the rubric midpoint of 3.0. Between‐model differences were significant ( p < 0.001), with a small‐to‐medium overall effect (Kendall's W = 0.17) and large within‐category effects ( W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier. Conclusion Retrieval‐augmented systems rated highest, but this advantage was confounded with response length. The dangerous‐response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.
Mayer et al. (Sun,) studied this question.