Study objective To compare a rule-based computable phenotype designed to identify patients with opioid use disorder in the emergency department (ED) with a large language model, using expert physician review as the reference standard. Methods We conducted a retrospective study of randomly sampled adult ED encounters (January 1, 2023 to October 17, 2024) at a single academic health system. We drew a stratified random sample based on whether encounters met a preexisting rule-based phenotype for identifying opioid use disorder. The phenotype incorporated diagnosis codes, medications for opioid use disorder, urine toxicology results, addiction consultations, and keyword matching. With zero-shot prompting, a large language model (ChatGPT 4.1) classified opioid use disorder using ED notes from the index visit. Two board-certified emergency physicians independently determined the presence of opioid use disorder by full chart review; discrepancies were adjudicated by a third reviewer. Using inverse probability weighting based on the sampling fractions, we estimated sensitivity, specificity, positive predictive value, and negative predictive value. Results Among 302 encounters, weighted opioid use disorder prevalence was 5.6% (95% confidence interval CI, 4.0 to 7.0%). The structured phenotype demonstrated sensitivity 0.84 (95% CI, 0.42 to 0.97) and specificity 0.964 (95% CI, 0.96 to 0.97) (positive predictive value 0.58; negative predictive value 0.99). The large language model demonstrated sensitivity 0.81 (95% CI 0.70-0.88) and specificity 0.996 (95% CI, 0.993 to 0.998) (positive predictive value 0.92; negative predictive value 0.99). Specificity was significantly higher for the large language model ( P <.0001). Conclusion Both approaches demonstrated strong diagnostic performance. Although the structured phenotype showed slightly higher sensitivity, the large language model achieved higher specificity and positive predictive value, suggesting potential to reduce false-positive alerts in ED workflows. Prospective validation in other populations is needed.
Molina et al. (Mon,) studied this question.