PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 22, 2026The American Surgeon0 citations

Artificial Intelligence in Surgical Education: A Pilot Study Using ASCRS Guideline-Derived Questions

View Full Paper
SPShivam PandyaTWTyler WilsonRMRyan Meyer

Key Points

  • To evaluate the accuracy and consistency of two large language models using ASCRS guidelines-based questions.
  • Developed 30 multiple-choice questions from ASCRS guidelines.
  • Validated questions by surgeon reviewers.
  • Presented questions to both AI models and calculated accuracy with confidence intervals.
  • Both models answered 29 out of 30 questions correctly, achieving 96.7% accuracy.
  • Statistical analysis showed performance significantly better than chance (p < .0001).
  • Inter-model agreement was perfect, with Cohen's kappa coefficient of 1.0.

Abstract

BackgroundLarge language models (LLMs) have demonstrated strong performance on general medical and surgical examinations; however, their capacity to accurately interpret and apply subspecialty clinical practice guidelines remains incompletely characterized.ObjectiveTo evaluate and compare the accuracy and consistency of two contemporary LLMs-Google Gemini and OpenEvidence-using multiple-choice questions (MCQs) derived directly from the 2022 American Society of Colon and Rectal Surgeons (ASCRS) Clinical Practice Guidelines for anorectal abscess, fistula-in-ano, and rectovaginal fistula.MethodsThirty guideline-based MCQs were developed and independently validated by surgeon reviewers. Each question was presented to both models under identical conditions without additional prompting. Accuracy was calculated with 95% confidence intervals and compared against chance performance (p0 = .25). Inter-model agreement was assessed using Cohen's kappa coefficient.ResultsBoth Gemini and OpenEvidence correctly answered 29 of 30 questions (96.7%; 95% CI, 0.83-0.999), significantly exceeding chance performance (P < .0001 for both models). Both models missed the same question, yielding perfect inter-model agreement (κ = 1.0).ConclusionIn this focused pilot study restricted to ASCRS anorectal disease guidelines, both LLMs demonstrated near-perfect and statistically equivalent accuracy. These findings suggest that contemporary LLMs can accurately apply subspecialty surgical guidelines within a narrow domain, though broader, multi-guideline evaluations are required before generalization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pandya et al. (2026) studied this question.

synapsesocial.com/papers/69e867136e0dea528ddeb610https://doi.org/10.1177/00031348261443342
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Artificial Intelligence in the Trauma Bay: A Pilot Comparison With Surgical Trainees2026
  2. 2Guideline concordance of large language models for ERAS colorectal surgery recommendations: a blinded, clinician-rated comparison of Google Gemini and ChatGPT2026
  3. 3Generative AI in Surgical Care: Evaluating Large Language Model Performance in Patient Education2025
  4. 4Large language models underperform in European general surgery board examinations: a comparative study with experts and surgical residents2025 · 4 citations
  5. 5Artificial Intelligence as a Support Tool for Preoperative Patient Education in Anesthesiology: A Comparative Evaluation of Five Large Language Models2026 · 1 citations