Abstract Purpose This systematic review evaluated the application of ChatGPT and other large language models in answering dental patient inquiries and explored their accuracy. Methods Following PRISMA guidelines, seven databases, including PubMed, Scopus, and Cochrane, were searched for studies published between November 2022 and June 2024. The review focused on publications addressing large language models’ performance in responding to patients’ questions, with studies evaluated for quality using the modified QUADAS‐2 framework. Data on accuracy were extracted, and a meta‐analysis was conducted. Subgroup and sensitivity analyses were performed to explore variations in performance and ensure robustness. Results A total of 25 studies were included, evaluating ChatGPT and other large language models. The pooled accuracy score for all large language models included was 81.87% (95% CI: 77.24%–86.51%), and 69.9% (95% CI: 57.3%–82.6%) of responses were considered clinically acceptable. Subgroup analysis revealed that the accuracy score of responses from ChatGPT‐3.5 was significantly higher than Microsoft Bing but not different from ChatGPT‐4.0 and Google Bard. Conclusion ChatGPT and other LLMs are promising alternatives for addressing patient inquiries and providing oral health education. However, challenges remain regarding accuracy, variability, and their ability to handle complex clinical scenarios, and further research is needed.
Zhang et al. (2025) studied this question.