Google Gemini showed similar guideline concordance to ChatGPT for ERAS colorectal surgery recommendations (mean Likert score 4.49 vs 4.44; p=0.604).
Cross-Sectional (n=52)
Blinded
Does Google Gemini provide better guideline concordance and safety compared to ChatGPT for ERAS colorectal surgery recommendations?
Both Google Gemini and ChatGPT showed high overall concordance with ERAS colorectal recommendations, but safety-relevant deviations occurred in 15-20% of responses, highlighting the need for clinician oversight.
Absolute Event Rate: 4.49% vs 4.44%
p-value: p=0.604
Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;
Sezer Gökçen (Tue,) conducted a cross-sectional in ERAS colorectal surgery recommendations (n=52). Google Gemini vs. OpenAI ChatGPT was evaluated on Guideline concordance on a 5-point Likert scale (p=0.604). Google Gemini showed similar guideline concordance to ChatGPT for ERAS colorectal surgery recommendations (mean Likert score 4.49 vs 4.44; p=0.604).
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: