Abstract Background Chronic obstructive pulmonary disease (COPD) requires personalised, guideline-based management according to the Global Initiative for Chronic Obstructive Lung Disease (GOLD) Report. As large language models (LLMs) are increasingly used for medical guidance, their adherence to updated clinical recommendations requires systematic evaluation. This study assessed and compared ChatGPT-5, DeepSeek-V3.2, Gemini 3, Grok, Manus, and Kimi K1.5 in delivering 2025 GOLD-consistent responses, focusing on accuracy, consistency, and clinical reasoning. Methods A standardised dataset of 90 questions derived from the 2025 GOLD Report was categorised into YES/NO, multiple-choice, and open-ended formats. During a 7-day longitudinal study in December 2025, all questions were administered daily across the six platforms, generating 3,780 responses. Performance was evaluated against the 2025 GOLD Report, and open-ended responses were assessed for quality and clinical comprehensiveness. Results Performance was assessed across two separate primary outcomes. For binary accuracy (YES/NO and MCQ questions), Kimi K1.5 achieved the highest rate (97.1%; 408/420; 95% CI, 95.1–98.4%), and Gemini 3 the lowest (89.5%; 376/420; 95% CI, 86.2–92.1%). For open-ended clinical reasoning quality, Gemini 3 achieved the highest weighted score (96.5%; 95% CI, 93.3–98.4%) and DeepSeek-V3.2 the lowest (91.8%; 95% CI, 87.4–94.9%), reflecting a performance inversion between the two outcome domains. All platforms exhibited statistical stability across the seven-day study period ( p > 0.05 for all models). Conclusion Large language models showed high but format-dependent adherence to the 2025 GOLD COPD guidelines, with clear divergence between binary accuracy and open-ended reasoning. While they may serve as supplementary tools under expert supervision, generalization to other settings requires further validation. Residual performance gaps persisted across both domains, with binary error rates ranging from 2.9% (Kimi K1.5) to 10.5% (Gemini 3), and open-ended weighted-score deficits ranging from 3.5% (Gemini 3) to 8.2% (DeepSeek-V3.2).
Al-waqeerah et al. (Sat,) studied this question.