Post-operative delirium (POD) is common after major surgery in older adults, and clear information about the risks and preventative strategies is not provided to these patients. Brief, surgery-specific POD handouts written in plain language are scarce. We evaluated whether POD handouts created by generative artificial intelligence (AI) achieve clinician-rated quality comparable to established patient leaflets. Cross-sectional pilot study examining patient-education documents based on independent expert ratings. Peri-operative patient education documents. Not applicable (no patient participants). Four expert raters. Eight independent 500-word POD handouts AI-generated with ChatGPT-4o using a standardized prompt. Comparators were the Ontario Health 2021 Delirium Patient Guide and the Royal College of Anaesthetists 2024 Becoming Confused After an Operation leaflet. Quality was rated from 4 to 16 using the Quality of Generated Language Outputs for Patients to generate a Language-Quality score. Raters assessed each domain, accuracy/completeness, bias, currency, and tone, on a scale from 1 to 4. Readability was assessed using Flesch-Kincaid grade level. Scores were analyzed using a linear mixed-effects model with provenance as a fixed effect and random intercepts for evaluator and document. Spearman correlation between readability and Language-Quality scores was conducted. Mean Language-Quality score was 14.19 (95% CI 13.68–14.70) for ChatGPT-4o handouts and 14.75 (95% CI 11.57–16.00) for controls. This difference was not statistically significant (95% CI -0.46 to 1.59, p = 0.25). Mean Flesch-Kincaid grade level was 7.0 for ChatGPT-4o handouts, 11.2 for Ontario Health, and 7.7 for Royal College handouts. Readability was not associated with Language-Quality score (Spearman r = 0.037, p = 0.92). AI-generated POD handouts achieved high clinician-rated quality with no significant differences with two commonly used comparator leaflets. Their mean readability grade level was lower than one widely used guide. Before dissemination, implementation should include checklist-based clinical verification and patient-centered evaluation.
Hua et al. (Wed,) studied this question.