Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
September 2, 2026npj VirusesOpen Access

Performance of large language models as a source of clinical information on bacteriophage therapy

View Full Paper
Ask AI
Bookmark
Share

Authors

NWNike WalterDADerek F. AmanatullahLDLaurent Debarbieux

Discussion

Loading...

Member takes

Overview

Expert evaluation study uncovers notable factual inaccuracies in AI-generated responses to patient questions, highlighting the need for clinician oversight before clinical use.

Key Points

  • To evaluate the quality, accuracy, completeness, clarity, and tone of large language model responses to patient inquiries regarding bacteriophage therapy.
  • Twelve bacteriophage therapy experts independently assessed LLM responses to 20 patient-relevant questions across accuracy, completeness, clarity, and tone/empathy using 5-point Likert scales (N=960 total ratings).
  • Adjusted mean scores ranged from 3.36 to 3.96 across domains, with significant differences among LLMs in completeness and tone/empathy (Holm-adjusted p = 0.042 for both; Cohen’s d = 0.12–0.29), but not in accuracy or clarity.
  • Claude achieved significantly lower completeness and tone/empathy scores, while Perplexity scored highest in completeness.
  • Experts recommended revisions for 34% to 40% of model responses and identified incorrect clinical information in 20% of answers.

Cite This Study

Walter et al. (2026) studied this question.

synapsesocial.com/papers/6a97e318c562ede874ec7b4chttps://doi.org/10.1038/s44298-026-00224-2
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of large language models as a source of clinical information on bacteriophage therapy2026
  2. 2Large language model responses to patient-oriented neurointerventional queries: A multirater assessment of accuracy, completeness, safety, and actionability2025
  3. 3Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists2026
  4. 4Real-world evaluation of large language model for patients medical and administrative queries in nuclear medicine2026
  5. 5Evaluating large language model responses to public questions on perianal abscess and anal fistula: a chart-informed comparative study2026