PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 2026Diagnostics2 citationsOpen Access

Performance of ChatGPT-4o, Gemini 2.0 Pro, and DeepSeek-V3 in Patient-Facing Information on Chest Wall Deformities: A Comparative Evaluation of Accuracy, RELIABILITY, and Reproducibility

DODeniz OkeOIOzge Gulsum IlleezEGEsra Giray

Key Points

  • This study aims to evaluate the accuracy, reliability, and reproducibility of three large language models in generating patient-facing content about chest wall deformities.
  • Developed eighty patient-facing questions across eight thematic domains.
  • Independently submitted questions to each model over two consecutive days.
  • Assessed accuracy using a four-point rubric evaluated by blinded physiatrists.
  • Evaluated reproducibility through agreement metrics and weighted Cohen’s kappa.
  • ChatGPT-4o achieved the highest median accuracy score of 1.20.
  • ChatGPT-4o had the lowest hallucination rate at 5.0%.
  • Gemini 2.0 Pro showed intermediate accuracy, while DeepSeek-V3 had the lowest accuracy and highest hallucination rate (11.25%).
  • Reproducibility was almost perfect for ChatGPT-4o with weighted κ values, followed by Gemini and DeepSeek-V3.

Abstract

Background: Large language models (LLMs) such as DeepSeek-V3, Google Gemini 2.0 Pro, and ChatGPT-4o are increasingly used by patients seeking online medical information. However, their accuracy, reliability, and reproducibility in patient-facing content related to chest wall deformities (CWD) remain unclear. This study aimed to compare the performance of three contemporary LLMs in generating information on pectus excavatum, pectus carinatum, and related thoracic deformities. Methods: Eighty patient-facing questions were developed across eight thematic domains and independently submitted to each model using newly created accounts over two consecutive days. Accuracy was assessed using a validated four-point rubric by blinded physiatrists, and reproducibility was evaluated using agreement metrics and weighted Cohen’s kappa. Results: ChatGPT-4o achieved the highest overall accuracy (median score: 1.20), the greatest proportion of fully accurate responses, and the lowest hallucination rate (5.0%). Gemini showed intermediate accuracy, while DeepSeek-V3 demonstrated the lowest accuracy and highest hallucination rate (11.25%). Across all models, general-information and quality-of-life domains had the best performance, whereas treatment-related questions showed the most errors. Reproducibility was highest for ChatGPT-4o (weighted κ = almost perfect), followed by Gemini and DeepSeek-V3. Inter-rater reliability was substantial (Fleiss’ κ = 0.69). Conclusions: Contemporary LLMs can generate largely accurate and reproducible patient-facing information on CWD, with ChatGPT-4o showing the strongest overall performance. This study provides the first domain-specific comparative evaluation of LLMs in CWD and integrates reproducibility analysis alongside accuracy and reliability assessment. While these tools may support patient education, treatment-related responses require caution, and LLMs should be used as adjuncts rather than substitutes for clinical counseling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Oke et al. (2026) studied this question.

synapsesocial.com/papers/6996a7e3ecb39a600b3edffchttps://doi.org/10.3390/diagnostics16040589
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1From Guidelines to Real‐Time Conversation: Expert‐Validated Retrieval‐Augmented and Fine‐Tuned GPT ‐4 for Hepatitis C Management2025 · 5 citations
  2. 2Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review2024 · 155 citations
  3. 3ChatGPT-4.0 or DeepSeek-V3? Comparative analysis of answers to the most frequently asked questions by total knee replacement candidate patients2025 · 8 citations
  4. 4Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information Access2023 · 56 citations
  5. 5Chatbot breakthrough in the 2020s? An ethical reflection on the trend of automated consultations in health care2021 · 157 citations