PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 2025Frontiers in Digital Health6 citationsOpen Access

Comparative performance evaluation of large language models in answering esophageal cancer-related questions: a multi-model assessment study

View Full Paper
ZHZhanwen HeLZLilan ZhaoGLGenglin Li

Key Points

  • All models scored above 4.0, indicating their overall effectiveness in answering esophageal cancer-related questions.
  • Gemini demonstrated superior accuracy, while ChatGPT excelled in providing comprehensive answers, particularly in surgical and postoperative contexts.
  • Retests of low-scoring responses led to overall quality improvements, although some answers lost completeness and relevance.
  • The findings suggest tailored applications of large language models in clinical settings, emphasizing the need for user education on their limitations.

Abstract

Background Esophageal cancer has high incidence and mortality rates, leading to increased public demand for accurate information. However, the reliability of online medical information is often questionable. This study systematically compared the accuracy, completeness, and comprehensibility of mainstream large language models (LLMs) in answering esophageal cancer-related questions. Methods In total, 65 questions covering fundamental knowledge, preoperative preparation, surgical treatment, and postoperative management were selected. Each model, namely, ChatGPT 5, Claude Sonnet 4.0, DeepSeek-R1, Gemini 2.5 Pro, and Grok-4, was queried independently using standardized prompts. Five senior clinical experts, including three thoracic surgeons, one radiologist, and one medical oncologist, evaluated the responses using a five-point Likert scale. A retesting mechanism was applied for the low-scoring responses, and intraclass correlation coefficients were used to assess the rating consistency. The statistical analyses were conducted using the Friedman test, the Wilcoxon signed-rank test, and the Bonferroni correction. Results All the models performed well, with average scores exceeding 4.0. However, the following significant differences emerged: Gemini excelled in accuracy, while ChatGPT led in completeness, particularly in surgical and postoperative contexts. Minor differences appeared in fundamental knowledge, but notable disparities were found in complex areas. Retesting showed improvements in overall quality, yet some responses showed decreased completeness and relevance. Conclusion Large language models have considerable potential in answering questions about esophageal cancer, with significant differences in completeness. ChatGPT is more comprehensive in complex scenarios, while Gemini excels in accuracy. This study offers guidance for selecting artificial intelligence tools in clinical settings, advocating for a tiered application strategy tailored to specific scenarios and highlighting the importance of user education to understand the limitations and applicability of LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

He et al. (2025) studied this question.

synapsesocial.com/papers/68e6679587ecc93a24d176f9https://doi.org/10.3389/fdgth.2025.1670510
Ask AI
Helpful
Bookmark
Share
View Full Paper