Key points are not available for this paper at this time.
OBJECTIVES: Large language models (LLMs) are increasingly used in health care communication but can inadvertently perpetuate stigmatizing language toward individuals with alcohol and substance use disorders. Despite growing interest in LLM performance, a focused evaluation of their propensity for SL and strategies to mitigate it remains lacking. METHODS: We generated 60 clinically relevant questions "prompts"; 20 each for alcohol use disorder (AUD), alcohol-associated liver disease (ALD), and substance use disorder (SUD) and tested 14 LLMs. Two physicians independently assessed all responses for stigmatizing language using guidelines from the National Institute on Drug Abuse and the National Institute on Alcohol Abuse and Alcoholism; discrepancies were resolved by a third physician. We employed iterative prompt engineering (PE)-a process of strategically crafting input instructions to guide model outputs towards nonstigmatizing language-to reduce stigmatizing language by incorporating a list of specific terms to avoid and identifying model-specific pitfalls. We compared the prevalence of SL in responses to native prompts (baseline, unengineered) versus engineered prompts, adjusting for word count in multivariate analyses. RESULTS: Of 840 responses generated from native prompts, 297 (35.4%) contained stigmatizing language, totaling 592 terms. With prompt engineering, only 53 (6.3%) of 840 responses contained stigmatizing language, comprising 104 terms. Prompts on topic of ALD yielded higher odds of stigmatizing language than those addressing AUD (adjusted odds ratio, 2.11; 95% CI, 1.47-3.02; P < 0.001), whereas prompts on substance use disorder (SUD) did not differ significantly from AUD (adjusted odds ratio, 1.17; 95% CI, 0.81-1.69; P = 0.40). Prompt engineering reduced the likelihood of stigmatizing language by 88% in univariate analysis (P < 0.001), and this effect persisted after adjusting for word count (adjusted odds ratio, 0.15; 95% CI, 0.11-0.20; P < 0.001). CONCLUSIONS: LLMs frequently generated stigmatizing language when discussing alcohol-related and substance use-related conditions, potentially undermining patient-centered care. However, targeted prompt engineering substantially reduced stigmatizing language occurrences across diverse models. These findings emphasize the need for ongoing model refinement and structured prompting strategies to ensure stigma-free language in health care communication.
Wang et al. (Thu,) studied this question.