PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025Cureus0 citationsOpen Access

Evaluating the Accuracy, Completeness, and Readability of Chatbot Responses to Refractive Surgery-Related Patient Questions: A Comparative Analysis of ChatGPT and Google Gemini

View Full Paper
SASerdar ArslanGKGüldeniz Usta Küçükbezirci

Key Points

  • ChatGPT and Google Gemini demonstrated strong reliability for responding to refractive surgery questions.
  • Accuracy scores were similar overall, but Gemini performed slightly worse on harder questions compared to ChatGPT.
  • The readability of responses was consistent, although ChatGPT provided more detailed answers and Gemini offered concise ones.
  • The complexity of questions influenced chatbot responses, emphasizing the necessity for continuous improvement in AI tools.

Abstract

Purpose This study evaluates the performance of ChatGPT and Google Gemini in addressing refractive surgery-related patient questions by analysing the accuracy, completeness, and readability of their responses. Methods A total of 40 refractive surgery-related questions were compiled and categorized into three levels of difficulty: easy, medium, and hard. Responses from ChatGPT and Google Gemini were blinded and evaluated by two experienced ophthalmologists using standardized criteria. Accuracy was scored on a six-point Likert scale, completeness on a three-point scale, and readability using Flesch-Kincaid Grade Level, Gunning Fog Index, Simple Measure of Gobbledygook (SMOG) Index, and word count. Intra- and inter-rater reliability were assessed using intra-class correlation coefficients (ICC). Results Both chatbots demonstrated high intra-rater (ICC>0.75) and inter-rater reliability. Accuracy scores were similar for most questions; however, statistically significant differences were observed for harder questions, where Gemini showed slightly reduced performance compared to ChatGPT. Readability metrics revealed no significant differences between the two tools, although ChatGPT responses tended to be more detailed, while Gemini generated more concise answers. Harder questions resulted in longer and more complex responses, as indicated by higher Gunning Fog and SMOG Index scores. Conclusions ChatGPT and Google Gemini exhibit strong potential in patient education, with complementary strengths in accuracy, readability, and response detail. The influence of question complexity on chatbot performance highlights the need for ongoing optimization to enhance both clarity and accessibility. These findings underscore the value of integrating artificial intelligence (AI) tools into healthcare to support patient education and engagement.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Arslan et al. (2025) studied this question.

synapsesocial.com/papers/68c1ad6a54b1d3bfb60e5e04https://doi.org/10.7759/cureus.88980
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1An Observational Study to Evaluate Readability and Reliability of AI-Generated Brochures for Emergency Medical Conditions2024 · 9 citations
  2. 2Physician and Artificial Intelligence Chatbot Responses to Cancer Questions From Social Media2024 · 109 citations
  3. 3Chatbots Utility in Healthcare Industry: An Umbrella Review2024 · 25 citations
  4. 4Recent Advances in Refractive Surgery: An Overview2024 · 34 citations
  5. 5The Technique of Clear Writing.1968 · 1,681 citations