Randomized trial assesses large language model responses versus human-generated answers for queries in nuclear medicine, suggesting potential usefulness.
Key Points
The aim is to evaluate the effectiveness of large language models in answering patient queries in nuclear medicine compared to human-generated responses.
Queries collected from patients were answered by nuclear medicine physicians, staff, and ChatGPT v4.1.
Responses were scored by experts using the QUEST framework and assessed for quality with binomial tests.
Inter-rater agreement was measured using Prevalence-Adjusted Bias-Adjusted Kappa (PABAK).
For medical queries, 76-98% of LLM responses rated equal or better than human responses (p < 0.001).
In administrative queries, 97% of non-experts found LLM responses more informative, with 86% preferring them.
PABAK indicated higher agreement for LLM responses (0.92-1.00) compared to human responses (-0.63 to -0.13).