Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
March 7, 2026Journal of Orthopaedic Trauma

Comparing The Efficacy Between ChatGPT 5, Grok 3, and Claude 4.5 Sonnet in Analyzing Orthopedic Trauma-Related Imaging

View Full Paper
Ask AI
Bookmark
Share

Authors

JHJoshua Aaron HolmstromMayo Clinic in ArizonaCBCollin BraithwaiteMayo Clinic in ArizonaAAAhmad R. AlhankawiMayo Clinic in Arizona

Discussion

Loading...

Member takes

Implication

Diagnostic performance comparison of AI platforms for analyzing orthopedic fractures, indicating mixed effectiveness.

Key Points

  • The research aims to evaluate the diagnostic capabilities of three AI models in identifying common orthopedic trauma fractures using imaging.
  • Conducted a retrospective comparison study
  • Utilized public online radiologic imaging databases
  • Assessed five common orthopedic trauma fractures
  • Evaluated ChatGPT 5, Grok 3, and Claude 4.5 Sonnet's diagnostic accuracy
  • Analyzed performance using radiographs and CT images
  • ChatGPT 5 diagnosed correctly in 26.8% of cases, Grok 3 in 18.8%, and Claude 4.5 Sonnet in 22.4%
  • Highest correct classification rates for fracture types were observed in ChatGPT 5
  • Sensitivity for ChatGPT 5 was 0.267, Grok 3 was 0.187, and Claude 4.5 Sonnet was 0.223
  • ChatGPT 5 and Grok 3 significantly outperformed Claude 4.5 Sonnet in diagnostic accuracy

Cite This Study

Holmstrom et al. (2026) studied this question.

synapsesocial.com/papers/69abc2455af8044f7a4ebba9https://doi.org/10.1097/bot.0000000000003166
View Full Paper
Ask AI
Bookmark
Share