PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Indian Journal of Orthopaedics3 citationsOpen Access

Evaluating ChatGPT-5’s Performance in Answering Common Patient Questions About Femoroacetabular Impingement and Hip Arthroscopy

View Full Paper
MVMaximilian VossHJHannah K. JaegerMSMikhail Salzmann

Key Points

  • The study aims to evaluate ChatGPT-5's effectiveness in answering patient questions about femoroacetabular impingement and hip arthroscopy.
  • ChatGPT-5 generated responses to 25 frequently asked questions regarding hip preservation.
  • Two hip preservation surgeons evaluated responses based on relevance, accuracy, clarity, and completeness.
  • Statistical analyses included mean scores, standard deviation, and inter-rater reliability measures.
  • Responses scored between 4.84 and 5.00 across evaluated domains, indicating high performance.
  • Inter-rater reliability showed moderate to excellent agreement (ICC values between 0.70–0.81).
  • No responses contained factually incorrect or unsafe information.

Abstract

Abstract Background Hip arthroscopy (HAS) is widely used to treat femoroacetabular impingement syndrome (FAIS), and many patients rely on online resources for medical information. Large language models (LLMs) such as ChatGPT have shown potential as supplementary educational tools in orthopedics; however, existing evaluations are limited to earlier model generations with variable accuracy and completeness. This study aimed to evaluate the accuracy, clarity, relevance, and completeness of ChatGPT-5 responses to common patient questions regarding FAIS and HAS. Methods ChatGPT-5 was used to generate 25 frequently asked patient questions and corresponding answers related to hip preservation. Two fellowship-trained hip preservation surgeons independently evaluated each response using a five-point Likert scale across four predefined domains: relevance, accuracy, clarity, and completeness. Descriptive statistics were calculated as mean ± standard deviation for each domain. Inter-rater reliability was assessed using a two-way random-effects intraclass correlation coefficient with absolute agreement (ICC 2, 1) and complemented by exact agreement percentages. Results All responses received excellent scores, with mean values ranging from 4.84 ± 0.27 (completeness) to 5.00 ± 0.00 (relevance). Accuracy (4.97 ± 0.08) and clarity (4.91 ± 0.17) were near-perfect. ICC values demonstrated moderate to excellent agreement (0.70–0.81), complemented by high exact agreement rates (84–100%). No answer contained factually incorrect, misleading, or unsafe information. Minor reductions in completeness were attributable to occasional brevity rather than substantive omissions. Conclusion ChatGPT-5 generated highly accurate, clear, and clinically appropriate patient-oriented explanations regarding FAIS and HAS, showing clear improvement compared with earlier ChatGPT versions. Although ChatGPT-5 represents a marked advancement in AI-based patient education, its use should be regarded as a complementary educational tool rather than a replacement for professional orthopedic counseling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Voss et al. (2026) studied this question.

synapsesocial.com/papers/6984343ff1d9ada3c1fb2307https://doi.org/10.1007/s43465-026-01696-3
Ask AI
Helpful
Bookmark
Share
View Full Paper