Background and Aims: High-quality colonoscopy requires accurate risk stratification for surveillance per the 2020 US Multi-Society Task Force (USMSTF) guidelines and adherence to 2024 ACG/ASGE quality benchmarks. Both are operationally challenging in clinical practice. We developed and validated Colon-Pilot, a large language model (LLM)--powered clinical decision support system (CDSS) using GPT-4o to automate and standardize both functions. Methods: The system was evaluated in two operational modes: (1) a human-in-the-loop clinical decision support validation of surveillance recommendations for 596 colonoscopies, comparing concordance with 2020 USMSTF guidelines against expert consensus; and (2) an automated administrative audit applying Colon-Pilot to 42,632 colonoscopies across the Mayo Clinic Health System to calculate 2024 ACG/ASGE priority quality indicators. Recommendations were auto-generated unless predefined safety criteria triggered manual review. Results: Colon-Pilot issued recommendations for 522/596 cases (87.6%) and flagged 12.4% for manual review. For automated cases, guideline-concordant accuracy was 97.5% (Cohen’s κ = 0.970) versus 69.7% (κ = 0.781) for original endoscopist recommendations. Discordant AI cases (n=13) most often recommended longer-than-appropriate intervals (62%). Applied to the enterprise dataset, Colon-Pilot calculated performance exceeding 2024 targets: Adenoma Detection Rate 49.8% (≥35%), Sessile Serrated Lesion Detection Rate 17.7% (≥6%), bowel preparation adequacy 91.8% (≥90%), and cecal intubation rate 97.4% (≥95%). Conclusion: Colon-Pilot demonstrated high fidelity in applying surveillance guidelines and automated quality benchmarking, outperforming unassisted endoscopists in guideline adherence. By combining safety protocols with large-scale automated reporting, it offers a scalable solution for improving both efficiency and quality in colorectal cancer prevention.
Garg et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: