Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
June 19, 2026npj Digital MedicineOpen Access

Real-world evaluation of large language model for patients medical and administrative queries in nuclear medicine

View Full Paper
Ask AI
Bookmark
Share

Authors

ALArash LatifoltojarAOAkintunde OrunmuyiMWMichael Wang

Discussion

Loading...

Member takes

Overview

Randomized trial assesses large language model responses versus human-generated answers for queries in nuclear medicine, suggesting potential usefulness.

Key Points

  • The aim is to evaluate the effectiveness of large language models in answering patient queries in nuclear medicine compared to human-generated responses.
  • Queries collected from patients were answered by nuclear medicine physicians, staff, and ChatGPT v4.1.
  • Responses were scored by experts using the QUEST framework and assessed for quality with binomial tests.
  • Inter-rater agreement was measured using Prevalence-Adjusted Bias-Adjusted Kappa (PABAK).
  • For medical queries, 76-98% of LLM responses rated equal or better than human responses (p < 0.001).
  • In administrative queries, 97% of non-experts found LLM responses more informative, with 86% preferring them.
  • PABAK indicated higher agreement for LLM responses (0.92-1.00) compared to human responses (-0.63 to -0.13).

Cite This Study

Latifoltojar et al. (2026) studied this question.

synapsesocial.com/papers/6a34dbc465a5b0777af2c6bfhttps://doi.org/10.1038/s41746-026-02889-8
View Full Paper
Ask AI
Bookmark
Share