Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
August 29, 2026Annals of The Royal College of Surgeons of EnglandOpen Access

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy

View Full Paper
Ask AI
Bookmark
Share

Authors

TDTP DavisBGB GuevelKLK Logishetty

Discussion

Loading...

Member takes

Overview

Blinded comparative study reveals higher clarity and completeness for ChatGPT over Gemini in peri-acetabular osteotomy queries, suggesting utility for surgical education.

Key Points

  • To compare the accuracy, completeness, clarity, and readability of ChatGPT and Google Gemini in addressing common patient questions regarding peri-acetabular osteotomy.
  • A panel of fellowship-trained peri-acetabular osteotomy (PAO) surgeons developed 10 common patient questions based on clinical experience.
  • Three fellowship-trained PAO surgeons, blinded to the response source, scored responses from each model on clarity, accuracy, and completeness using a 5-point Likert scale.
  • Readability for both models was calculated using Flesch-Kincaid Reading Ease scores.
  • ChatGPT achieved significantly higher ratings for completeness and clarity than Gemini (p = 0.006), requiring minimal clarification.
  • Gemini demonstrated reduced specificity and minor inaccuracies that diminished its perceived clinical reliability.
  • Both models produced responses characterized by similarly 'difficult' Flesch Reading Ease scores.

Cite This Study

Davis et al. (2026) studied this question.

synapsesocial.com/papers/6a92995f8e5d7d1fc0c11444https://doi.org/10.1308/rcsann.2026.0056
View Full Paper
Ask AI
Bookmark
Share