PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 16, 2026Endoscopy3 citations

Generative artificial intelligence for patient education material on gastric cancer prevention

TRTommy RizkalaNMNatasha Stephens MuenchCHCesare Hassan

Key Points

  • This research aims to evaluate the effectiveness of generative AI in creating educational materials for gastric cancer prevention.
  • Pilot study with a two-period, crossover, blinded design.
  • Compared ChatGPT-4o summaries to Digestive Cancers Europe (DiCE) summaries.
  • Expert physicians and a patient advisory committee rated materials for various criteria.
  • Readability was measured using Flesch–Kincaid and SMOG indices.
  • Expert ratings showed no significant differences in accuracy, completeness, comprehensibility, and satisfaction between ChatGPT-4o and DiCE.
  • Median ratings for accuracy were similar, with scores of 5 for both summaries.
  • Patient ratings paralleled expert assessments, showing consistency in evaluation.
  • Readability did not meet recommended guidelines for both summaries.

Abstract

Background This study assessed the effectiveness of large language models (LLMs) in generating lay summaries for patient education on the management of precancerous lesions and early neoplasia in the stomach. Methods In this pilot study, we used a two-period, crossover, blinded design to compare a ChatGPT-4o summary versus a Digestive Cancers Europe (DiCE) summary. Two panels rated the materials: expert physicians and DiCE Patient Advisory Committee members. Experts scored accuracy, completeness, comprehensibility, and satisfaction (across five sections); patients rated overall completeness, comprehensibility, and satisfaction. Paired comparisons used mixed-effects estimates. Readability was assessed with Flesch–Kincaid grade level (FKGL) and SMOG index. Results Median expert ratings were similar between materials across metrics. For the overall summary, median (range; IQR) scores were: accuracy 5 (4–6; 1) for ChatGPT-4o vs. 5 (3–6; 1) for DiCE (P = 0.10); completeness 4 (3–5; 1) vs. 4 (2–5; 1; P = 0.27); comprehensibility 4 (3–5; 1) vs. 4 (2–5; 1; P = 0.33); and satisfaction 4 (2–5; 1) vs. 3 (1–5; 2; P = 0.53). Patient ratings mirrored experts, with very similar results. Readability failed to meet guideline recommendations for both summaries on both FKGL and SMOG scores. Conclusion ChatGPT-4o produced patient materials comparable to DiCE, but both require readability optimization; a human-in-the-loop workflow and future tests across prompts and models are warranted.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rizkala et al. (2026) studied this question.

synapsesocial.com/papers/6992b45f9b75e639e9b09528https://doi.org/10.1055/a-2780-0664
Ask AI
Helpful
Bookmark
Share
View Full Paper