Blinded comparative study reveals higher clarity and completeness for ChatGPT over Gemini in peri-acetabular osteotomy queries, suggesting utility for surgical education.
Key Points
To compare the accuracy, completeness, clarity, and readability of ChatGPT and Google Gemini in addressing common patient questions regarding peri-acetabular osteotomy.
A panel of fellowship-trained peri-acetabular osteotomy (PAO) surgeons developed 10 common patient questions based on clinical experience.
Three fellowship-trained PAO surgeons, blinded to the response source, scored responses from each model on clarity, accuracy, and completeness using a 5-point Likert scale.
Readability for both models was calculated using Flesch-Kincaid Reading Ease scores.
ChatGPT achieved significantly higher ratings for completeness and clarity than Gemini (p = 0.006), requiring minimal clarification.
Gemini demonstrated reduced specificity and minor inaccuracies that diminished its perceived clinical reliability.
Both models produced responses characterized by similarly 'difficult' Flesch Reading Ease scores.