PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 13, 2026Cukurova Anestezi ve Cerrahi Bilimler Dergisi0 citationsOpen Access

Guideline concordance of large language models for ERAS colorectal surgery recommendations: a blinded, clinician-rated comparison of Google Gemini and ChatGPT

View Full Paper
SGSezer Gökçen

Key Result

Google Gemini showed similar guideline concordance to ChatGPT for ERAS colorectal surgery recommendations (mean Likert score 4.49 vs 4.44; p=0.604).

Key Points

  • This study evaluates the concordance of two large language models with ERAS guidelines for colorectal surgery.
  • Converted ERAS recommendations into a 52-question bank covering preoperative, intraoperative, and postoperative elements.
  • Conducted blinded evaluations of responses from Google Gemini and ChatGPT, rated by two clinicians.
  • Analyzed data using paired Wilcoxon signed-rank tests for concordance and McNemar tests for safety flags.
  • Mean Likert concordance was 4.49 for Gemini and 4.44 for ChatGPT (p=0.604).
  • Safety flags occurred in 20.2% of Gemini ratings and 15.4% of ChatGPT ratings (p=0.383).
  • Inter-rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).

Study Design

Type

Cross-Sectional (n=52)

Blinding

Blinded

Structured PICO

Does Google Gemini provide better guideline concordance and safety compared to ChatGPT for ERAS colorectal surgery recommendations?

P
Population
52 questions based on the 2025 ERAS Society recommendations for elective colorectal surgery, evaluated by two clinicians for guideline concordance and safety.
E
Exposure
Google Gemini (web interface)
C
Comparator
OpenAI ChatGPT (web interface)
O
Outcome
Guideline concordance on a 5-point Likert scale and safety risk (0=none, 1=potential harm, 2=critical harm)

Both Google Gemini and ChatGPT showed high overall concordance with ERAS colorectal recommendations, but safety-relevant deviations occurred in 15-20% of responses, highlighting the need for clinician oversight.

Main Result

Absolute Event Rate: 4.49% vs 4.44%

p-value: p=0.604

Limitations

  • Safety-relevant deviations were not rare
  • Rater agreement varied by model

Abstract

Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sezer Gökçen (2026) conducted a cross-sectional in ERAS colorectal surgery recommendations (n=52). Google Gemini vs. OpenAI ChatGPT was evaluated on Guideline concordance on a 5-point Likert scale (p=0.604). Google Gemini showed similar guideline concordance to ChatGPT for ERAS colorectal surgery recommendations (mean Likert score 4.49 vs 4.44; p=0.604).

synapsesocial.com/papers/6a54a170da93060f8148c31chttps://doi.org/10.36516/jocass.1892872
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluation and mitigation of the limitations of large language models in clinical decision-making2024 · 680 citations
  2. 2TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods2024 · 3,327 citations
  3. 3Postoperative complication management: How do large language models measure up to human expertise?2025 · 2 citations
  4. 4Guidelines for perioperative care in elective colorectal surgery: Enhanced Recovery After Surgery (ERAS) Society recommendations 20252025 · 350 citations
  5. 5Guidelines for Perioperative Care in Elective Colorectal Surgery: Enhanced Recovery After Surgery (ERAS ® ) Society Recommendations: 20182018 · 2,125 citations