PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026npj Health Systems3 citationsOpen Access

Detecting stigmatizing language in clinical notes with large language models for addiction care

View Full Paper
RSRohan SethiJCJohn CaskeyYGYanjun Gao

Key Points

  • This research aims to evaluate how well large language models can detect stigmatizing language in clinical notes related to substance use disorders.
  • Annotated a dataset of over 77,000 clinical notes from the MIMIC-III database for stigma detection.
  • Utilized Meta’s Llama-3 8B Instruct LLM with various approaches including zero-shot, in-context learning, supervised fine-tuning, and keyword search.
  • Evaluated the performance of models on a held-out test set and an external validation dataset from the University of Wisconsin Health System.
  • Supervised fine-tuning achieved the highest accuracy of 97.2%, with in-context learning following.
  • Both in-context learning and supervised fine-tuning effectively identified stigmatizing language that was missed in initial annotations.
  • In external validation, supervised fine-tuning performed at 97.9% accuracy.

Abstract

Abstract Intensive care units (ICU) produce numerous progress notes that may contain stigmatizing language that perpetuate negative biases and punitive approaches against patients. Patients with substance use disorders are particularly vulnerable to stigma. This study examined the performance of Large Language Models (LLMs) in the identification of stigmatizing language. We annotated a dataset with over 77,000 stigmatizing and non-stigmatizing notes from the MIMIC-III database. We utilized Meta’s Llama-3 8B Instruct LLM to run the following experiments for stigma detection: zero-shot; in-context learning; in-context learning with a selective retrieval; supervised fine-tuning (SFT); and keyword search. All approaches were evaluated on a held-out test set and external validation (University of Wisconsin Health System). SFT had the best performance with 97.2% accuracy, followed by in-context learning. The LLMs with in-context learning and SFT provided appropriate reasoning for false positives during human review. Both approaches identified clinical notes with stigmatizing language that were missed during annotation. SFT achieved 97.9% accuracy on external validation dataset. LLMs, particularly SFT and in-context learning, effectively identify stigmatizing language in ICU notes with high accuracy while explaining their reasoning in an asynchronous fashion and demonstrated the ability to identify novel stigmatizing language, not explicitly in training data nor existing guidelines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sethi et al. (2026) studied this question.

synapsesocial.com/papers/698433f6f1d9ada3c1fb1868https://doi.org/10.1038/s44401-026-00069-0
Ask AI
Helpful
Bookmark
Share
View Full Paper