PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 2, 2025npj Digital Medicine8 citationsOpen Access

Benchmarking proprietary and open-source language and vision-language models for gastroenterology clinical reasoning

View Full Paper
SRSara RafieeBKBara El Kurdi

Key Points

  • Language models demonstrated variable accuracy in clinical reasoning assessments related to gastroenterology.
  • Top-performing models included proprietary options, achieving accuracy percentages like 82% and 74%.
  • Assessment utilized board-style multiple-choice questions to evaluate model performance across different metrics.
  • Findings highlight the importance of quantization processes in enhancing model capabilities for real-world applications.

Abstract

This study evaluated the effectiveness of large language models (LLMs) and vision-language models (VLMs) in gastroenterology. We used board-style multiple-choice questions to assess the performance of both proprietary and open-source LLMs and VLMs-including GPT, Claude, Gemini, Mistral, Llama, Mixtral, Phi, and Qwen, across different interfaces, computing environments, and levels of compression (quantization). Among the proprietary models, o1-preview (82.0%) and Claude3.5-Sonnet (74.0%) had the highest accuracy, outperforming the top open-source models: Llama3.3-70b (65.7%) and Qwen-2.5-72b (61.0%). Among the small quantized open-source models, the 8-bit Llama 3.2-11b (51.7%) and 6-bit Phi3-14b (48.7%) performed the best, with scores comparable to their full-precision counterparts. Notably, VLM accuracy on image-containing questions improved (~10%) when given human-generated captions, remained unchanged with original images, and declined with LLM-generated captions. Further research is warranted to evaluate model capabilities in real-world clinical decision-making scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rafiee et al. (2025) studied this question.

synapsesocial.com/papers/692e3d626c9b3ab28c186ba8https://doi.org/10.1038/s41746-025-02174-0
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Toward expert-level medical question answering with large language models2025 · 914 citations
  2. 2Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models2023 · 3,819 citations
  3. 3What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation2024 · 2 citations
  4. 4Su1962 GI-COPILOT: AUGMENTING CHATGPT WITH GUIDELINE-BASED KNOWLEDGE2024 · 1 citations
  5. 5DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text2024 · 5 citations