Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 15, 2026JACCP JOURNAL OF THE AMERICAN COLLEGE OF CLINICAL PHARMACY

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists

View Full Paper
Ask AI
Bookmark
Share

Authors

NBNuntapong BoonritNBNajwa Bin‐usengASAphichaya Sirijariyawat

Discussion

Loading...

Member takes

Overview

Comparative study reveals high relevance but variable concordance and poor citation credibility among large language models answering drug queries, highlighting the need for pharmacist oversight.

Key Points

  • To evaluate the performance, clarity, clinical concordance, and citation credibility of multiple large language models compared to pharmacist-provided answers for real-world drug information questions.
  • Evaluated 714 responses generated across seven LLMs responding to 102 real-world drug information questions from a university hospital in Thailand.
  • Assessed clarity in Thai, relevance, context awareness, concordance against pharmacist reference standards, and citation credibility following an inter-rater reliability pilot phase (N=70 responses, 3 assessors; ICC/Fleiss' kappa: 0.79–0.86).
  • Analyzed domain fulfillment and binary outcomes using Cochran's Q test and post hoc McNemar testing.
  • LLMs achieved high fulfillment in clarity (85.25%–92.75%), relevance (0.97–1.00), and context awareness (0.90–0.99).
  • Concordance with pharmacist answers ranged from 0.70 to 0.86, with a statistically significant difference observed between ChatGPT-5.2 Thinking and Copilot Think Deeper (McNemar test, adjusted p = 0.0178).
  • Citation credibility remained consistently poor across all evaluated models, ranging from 0.03 to 0.36.

Cite This Study

Boonrit et al. (2026) studied this question.

synapsesocial.com/papers/6a8019db75c2e31742c8607bhttps://doi.org/10.1002/jac5.70269
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Capabilities of Large Language Models in Detecting and Managing Drug Interactions During Medication Reviews: Potential Implications as A Digital Assistant for Pragmatic Pharmacy Practice in Thailand2025
  2. 2Large language model responses to patient-oriented neurointerventional queries: A multirater assessment of accuracy, completeness, safety, and actionability2025 · 2 citations
  3. 3Real-world evaluation of large language model for patients medical and administrative queries in nuclear medicine2026
  4. 4Evaluating the Use of Large Language Models to Answer Patient-Facing Clinical Trial Questions2025
  5. 5Real-world evaluation of large language models in detecting drug-related problems: A clinical pharmacist–AI concordance study in hematology care2026 · 1 citations