PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026Drug and Alcohol Review0 citationsOpen Access

Large Language Models Accurately Identify People Who Inject Drugs From Infectious Diseases Discharge Summaries in an Australian Hospital

View Full Paper
DGDavid Goodman‐MezaMMMarianne MartinelloJMJeffrey Masters

Key Points

  • To evaluate the effectiveness of large language models in identifying people who inject drugs from hospital discharge summaries.
  • Cross-sectional analysis of discharge summaries from 2018 to 2022 at St Vincent's Hospital, Sydney.
  • De-identified summaries annotated for PWID status, drug use, injection recency, and treatment.
  • Comparison of eight LLMs based on prevalence-weighted average-F1 scores.
  • 149 out of 859 admissions identified as PWID (17.1%).
  • Best model (Llama 3.3) achieved average-F1 of 0.845 (95% CI 0.733, 0.927).
  • Sensitivity for injecting drug use was 0.819 (95% CI 0.753, 0.879) and specificity was 0.999 (95% CI 0.996, 1.00).

Abstract

INTRODUCTION: People who inject drugs (PWID) face a high risk for serious infections, yet International Classification of Diseases (ICD) codes fail to identify this population. Large language models (LLM) offer a promising alternative by extracting information from unstructured clinical text. This study evaluated the diagnostic performance of off-the-shelf LLMs in identifying PWID and related attributes from hospital discharge summaries. METHODS: In this cross-sectional study, discharge summaries from the Infectious Diseases service at St Vincent's Public Hospital, Sydney, between 2018 and 2022 were reviewed. A single reviewer manually annotated each de-identified summary for PWID status, drugs reported, injection recency and opioid agonist therapy. Eight LLMs (Gemma3, Llama 3.3, Mistral, Phi4, hippomistral, llama3-med 8B and 70B and OpenBioLLM) were compared using prevalence-weighted average-F1 scores. Diagnostic metrics with bootstrapped 95% confidence intervals were calculated for each annotated category. RESULTS: Of 859 first admissions, manual review identified 149 (17.1%) PWID. ICD codes showed low sensitivity (≤ 0.32) but high specificity (≥ 0.97) for identifying PWID. The best-performing model (Llama 3.3) achieved a prevalence-weighted average-F1 of 0.845 (0.733, 0.927). For injecting drug use, sensitivity was 0.819 (95% CI 0.753, 0.879) and specificity 0.999 (0.996, 1.00). Identification of heroin, methamphetamine, cannabis and methadone was near perfect (F1 > 0.973), while illicit prescription opioid and benzodiazepine use were identified less accurately (F1 = 0.400 and 0.606). DISCUSSION AND CONCLUSIONS: LLMs accurately identify PWID from discharge summaries, outperforming ICD codes. Challenges remain for certain substances, underscoring the need for task-specific tuning, external validation and integration with structured data to enhance surveillance and interventions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Goodman‐Meza et al. (2026) studied this question.

synapsesocial.com/papers/6a095bba7880e6d24efe18b9https://doi.org/10.1111/dar.70169
Ask AI
Helpful
Bookmark
Share
View Full Paper