BackgroundChatbots and large language models, particularly ChatGPT, have led to an increasing number of studies on the potential for chatbots in patient education. In this systematic review, we aimed to provide a pooled assessment of the appropriateness and accuracy of chatbot responses in patient education across various medical disciplines.MethodsThis was a PRISMA-compliant systematic review and meta-analysis. PubMed and Scopus were searched from January-August 2023. Eligible studies that assessed the utility of chatbots in patient education were included. Primary outcomes were the appropriateness and quality of chatbot responses. Secondary outcomes included readability and concordance with published guidelines and Google searches. A random-effect proportional meta-analysis was used for pooling data.ResultsFollowing initial screening, 21 studies were included. The pooled rate of appropriateness of chatbot answers was 89.1% (95%CI: 84.9%-93.3%). ChatGPT was the most assessed chatbot. Responses, while accurate, were found to be at a college reading level as the weighted mean Flesh-Kincaid Grade Level was 13.1 (95%CI: 11.7-14.5) and the weighted mean Flesch Reading Ease Score was 38.6 (95%CI: 29- 48.2). Answers of chatbots to questions relevant to patient education had 78.6%-95% concordance with published guidelines in colorectal surgery and urology. Chatbots had higher patient education scores (87% vs 78%) than Google Search.ConclusionsChatbots provide largely accurate and appropriate answers for patient education. The advanced reading level of chatbot responses might be a limitation to their wide adoption as a source for patient education. However, they outperform traditional search engines and align well with professional guidelines, showcasing their potential in patient education.
Emile et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: