PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 29, 20245 citationsOpen Access

Answering real-world clinical questions using large language model based systems

View Full Paper
YLYen LowMJMichael L. JacksonRHRebecca J. Hyde

Key Points

Key points are not available for this paper at this time.

Abstract

Evidence to guide healthcare decisions is often limited by a lack of relevant and trustworthy literature as well as difficulty in contextualizing existing research for a specific patient. Large language models (LLMs) could potentially address both challenges by either summarizing published literature or generating new studies based on real-world data (RWD). We evaluated the ability of five LLM-based systems in answering 50 clinical questions and had nine independent physicians review the responses for relevance, reliability, and actionability. As it stands, general-purpose LLMs (ChatGPT-4, Claude 3 Opus, Gemini Pro 1.5) rarely produced answers that were deemed relevant and evidence-based (2% - 10%). In contrast, retrieval augmented generation (RAG)-based and agentic LLM systems produced relevant and evidence-based answers for 24% (OpenEvidence) to 58% (ChatRWD) of questions. Only the agentic ChatRWD was able to answer novel questions compared to other LLMs (65% vs. 0-9%). These results suggest that while general-purpose LLMs should not be used as-is, a purpose-built system for evidence summarization based on RAG and one for generating novel evidence working synergistically would improve availability of pertinent evidence for patient care.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Low et al. (2024) studied this question.

synapsesocial.com/papers/68e629a8b6db6435875bc619https://doi.org/10.48550/arxiv.2407.00541
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence2025
  2. 2LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responses2025 · 12 citations
  3. 3Evaluating Large Language Models for Evidence-Based Clinical Question Answering2025
  4. 4LLM-assisted systematic review of large language models in clinical medicine2026 · 50 citations
  5. 5Performance of large language models in numerical vs. semantic medical knowledge: Benchmarking on evidence-based Q&As2024