Key points are not available for this paper at this time.
Background/Objectives: Inborn errors of immunity (IEI) are rare and complex pediatric disorders that create significant information gaps for families and non-specialist healthcare professionals. Large language models (LLMs) such as ChatGPT are increasingly used as on-demand health information resources; however, evidence on their performance in rare pediatric diseases remains limited. This study aimed to evaluate the reliability, quality, readability, understandability, reproducibility, and safety-related concerns of ChatGPT-4o responses to frequently searched questions about pediatric IEI posed by healthcare professionals and patients/caregivers. Methods: This cross-sectional evaluation used the publicly accessible ChatGPT-4o interface to generate responses to 20 frequently searched questions about pediatric IEI, equally distributed between healthcare professional (n = 10) and patient/caregiver queries (n = 10). Three pediatric allergy-immunology specialists independently evaluated response quality using the modified DISCERN (mDISCERN) and Global Quality Scale (GQS) tools, supplemented by a structured expert-based assessment of misinformation, safety-related concerns, suspected factual issues, missing disclaimers, and clinically meaningful inter-iteration inconsistency. Text readability was assessed using four validated indices (ARI, FRES, FKGL, GFR), comprehensibility using the Patient Education Materials Assessment Tool (PEMAT), and reproducibility using natural language processing methods. Results: ChatGPT-4o demonstrated strong overall performance, with median mDISCERN and GQS scores of 4 (IQR: 3–5) for both query types. Readability scores substantially exceeded recommended thresholds, with FKGL scores of 12.96 ± 0.69 and 10.83 ± 0.67 for professional and patient/caregiver queries, respectively. Mean PEMAT understandability scores were 71.80 ± 5.75% for professional queries and 80.80 ± 4.73% for patient/caregiver queries (p = 0.001). Reproducibility was high, with semantic similarity rates of 86.10 ± 3.84% and 87.30 ± 3.68%, respectively. Suspected factual issues were identified in 4 of 20 responses (20%), safety-related concerns in 3 (15%), clinically meaningful inter-iteration inconsistencies in 3 (15%), and missing medical disclaimers in all 20 responses (100%). Conclusions: ChatGPT-4o showed strong performance across validated quality metrics for pediatric IEI information support; however, its high reading level, universal absence of medical disclaimers, and occasional clinically meaningful inconsistencies limit its suitability as a standalone source for clinically sensitive guidance. These findings underscore the need for AI-driven patient education tools with improved readability, adaptive complexity adjustment, and safety-oriented communication.
Taşkırdı et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: