Abstract Background Chronic Obstructive Pulmonary Disease (COPD) remains to be a leading cause of morbidity and mortality worldwide. One of the significant challenges to its diagnosis and management is awareness due to limitations in access to guideline-based information, especially in low- and middle-income countries like the Philippines. Artificial intelligence (AI), particularly large language models (LLMs), is now of interest and is thought to be a potential tool to address information gaps, but its consistency with established clinical practice guidelines and utility in local contexts requires further evaluation. This study compared the quality of information provided by AI language models and Frequently-asked questions (FAQs) answers in the diagnosis and management of COPD based on GOLD guidelines. Methods This study employed an analytic, cross-sectional design that compared the accuracy, completeness, and clarity of AI-generated responses from Google Gemini, ChatGPT, DeepSeek, and Meta.ai to Frequently Asked Questions (FAQs) based on the 2025 GOLD guidelines for COPD diagnosis and management. Common COPD-related questions were provided to each AI platform in both English and Filipino language. Five board-certified experts in Pulmonary Medicine evaluated anonymized responses independently using standardized Likert scales. Mean scores were obtained and compared across the AI platforms and against responses from GOLD guidelines using one-way ANOVA. Inter-rater reliability was assessed with the Intraclass Correlation Coefficient (ICC). Results All four AI platforms produced responses of comparable quality in both English and Filipino language. While ChatGPT and Google Gemini achieved slightly higher mean scores for accuracy (up to 5.35/6), completeness (up to 2.55/3), and clarity (up to 2.85/3), these differences were not statistically significant (p 0.05). On all measures, however, the GOLD guideline responses remained significantly superior across all metrics (p 0.95). Conclusion AI-generated responses from top LLMs approach the quality of standardized guideline-based COPD FAQs, but are statistically inferior to expert-reviewed GOLD responses. Nonetheless, all platforms delivered comparable responses in English and Filipino. Results from this study can provide some insights to clinicians regarding the use of AI in healthcare and patient education, including the choice of LLMs that are most reliable, proper formulation of prompts which suits best their needs. While AI platforms can improve patient access to medical information and serve as a good starting point, they should not be used in place of expert assistance. Further evaluation, greater assessor inclusion, and integration of patient viewpoints are recommended to maximize the responsible use of AI in pulmonary medicine and patient education. This abstract is funded by: none
Gonzales et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: