Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek
Comparative trial evaluates responses of ChatGPT, Gemini, and DeepSeek for hemorrhoid-related patient queries, suggesting varying communication styles and adequacy.
Key Points
To compare the clinical accuracy, safety, and adequacy of responses from three large language models to hemorrhoid-related questions.
Cross-sectional comparative study with 25 hemorrhoid-related questions categorized into three subgroups.
Responses evaluated by two surgeons using a structured 5-point scoring system for accuracy, safety, and appropriateness.
Statistical comparisons were performed using Friedman and post hoc tests with Bonferroni correction.
Overall quality differed among models (χ2(2) = 29.119, p < 0.001, Kendall’s W = 0.582).
ChatGPT scored highest (5.00 ± 0.00), followed by Gemini (4.80 ± 0.41) and DeepSeek (4.12 ± 0.67).
Model differences were evident in clinical scenarios with alarm symptoms, while no harmful recommendations were made.