Abstract Background and Objective Pulmonary rehabilitation (PR) is essential for chronic obstructive pulmonary disease (COPD) management, yet patients face significant barriers in accessing high-quality educational information. Large language models (LLMs) like ChatGPT offer potential solutions for democratizing health information access, but their accuracy, reproducibility, and safety in COPD rehabilitation contexts remain unclear. Moreover, even if LLM-generated content is clinically accurate, whether patients possess sufficient skills to effectively utilize these tools needs to be elucidated. This study aimed to evaluate ChatGPT-generated educational content for COPD PR through expert assessment of clinical quality and direct examination of patient proficiency in information retrieval. Methods ChatGPT-4o generated responses to 19 questions predetermined through a two-round Delphi process in independent chat sessions with memory function disabled. To assess reproducibility, queries were conducted twice with a three-day interval (July 21 and 24, 2025). Then we evaluated the generated answers through TF-IDF vectorization with cosine similarity for lexical patterns and OpenAI’s embedding model for semantic consistency. Twenty experts independently evaluated responses using a 4-point Likert scale across four dimensions: accuracy (medical evidence basis), readability (patient-comprehensible terminology), safety (guidance for patient-specific variations), and overall quality. Inter-rater reliability was assessed using Gwet’s AC2 with quadratic weights. Descriptive analyses were conducted for the quality assessments. Additionally, Ten COPD patients from an outpatient clinic participated in an information retrieval task, seeking COPD rehabilitation information from ChatGPT via voice input. Results Twenty experts demonstrated high inter-rater agreement, with Gwet’s AC2 of 0.87 (95% CI: 0.85-0.89). High contextual similarity between duplicate runs of each question was observed based on OpenAI’s embedding model, with a mean of 0.94 (SD = 0.03, range: 0.87-0.98). Regarding clinical feasibility, ChatGPT demonstrated consistently high performance across all four evaluation domains, with mean scores of Accuracy 3.63, Readability 3.57, Safety 3.68, and Overall Quality 3.64. Despite the acceptable performance of ChatGPT, only 2 of 10 patients with COPD successfully retrieved COPD-related rehabilitation information via ChatGPT. Most patients were unaware of the need to disclose their disease profile to ChatGPT to obtain personalized information. Conclusions Expert evaluation demonstrated that ChatGPT-generated educational content for COPD pulmonary rehabilitation was acceptable in terms of reproducibility and accuracy. However, patient testing revealed that despite the quality of AI-generated information, patients lacked appropriate prompting skills to effectively access this information. These findings highlight the need for structured education and training programs to help patients develop adequate prompting techniques, verify AI-generated information accurately, and make informed decisions. This abstract is funded by: This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2025-00519997).
Lee et al. (Fri,) studied this question.