Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 17, 2026Open Access

Performance of large language models as a source of clinical information on bacteriophage therapy

View Full Paper
Ask AI
Bookmark
Share

Authors

NWNike WalterDADerek F. AmanatullahLDLaurent Debarbieux

Discussion

Loading...

Member takes

Overview

Comparative evaluation reveals variable completeness across large language models answering bacteriophage therapy queries, indicating that artificial intelligence drafts require expert clinical...

Key Points

  • To assess the quality, accuracy, completeness, clarity, and empathy of large language model responses to patient questions concerning bacteriophage therapy.
  • Evaluated responses from multiple large language models to 20 patient-relevant questions on bacteriophage therapy.
  • Engaged 12 clinicians and research experts in bacteriophage therapy to independently score each response across four domains using 5-point Likert scales, yielding 960 total ratings.
  • Adjusted mean scores ranged from 3.36 to 3.96 across domains, with significant differences among models observed for completeness and tone/empathy (Holm-adjusted p = 0.042 for both; Cohen’s d = 0.12–0.29), but not for accuracy or clarity.
  • Claude scored significantly lower than other models for completeness and tone/empathy, whereas Perplexity achieved the highest completeness scores.
  • Experts recommended improvements for 34% to 40% of responses, identifying incorrect information in 20% of evaluated outputs.

Cite This Study

Walter et al. (2026) studied this question.

synapsesocial.com/papers/6aabb61e5f706d05830e4943https://doi.org/10.5283/epub.80717
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of large language models as a source of clinical information on bacteriophage therapy2026
  2. 2Large language model responses to patient-oriented neurointerventional queries: A multirater assessment of accuracy, completeness, safety, and actionability2025 · 2 citations
  3. 3Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists2026
  4. 4Real-world evaluation of large language model for patients medical and administrative queries in nuclear medicine2026
  5. 5Evaluating large language model responses to public questions on perianal abscess and anal fistula: a chart-informed comparative study2026