Abstract Introduction Opioid use disorder (OUD) is common in emergency departments (EDs); identification via structured computable phenotypes may miss important clinical context. Objective Compare a computable structured OUD phenotype with a zero shot large language model (LLM) using expert review as the reference. Methods We retrospectively analyzed 202 adult ED encounters. Two emergency physicians independently determined OUD status with consensus adjudication. The phenotype used ICD 10 codes, medications, toxicology, consult notes, and keyword rules. The LLM (GPT 4.1) classified OUD from concatenated ED notes. Test characteristics and McNemar's tests were computed. Results Experts classified 56 (28%) encounters as OUD (kappa=0.77). The phenotype showed sensitivity 0.98 (95% CI 0.93-1.00) and specificity 0.54 (0.44-0.64). The LLM showed sensitivity 0.93 (0.87-0.96) and specificity 0.90 (0.77-0.97). Sensitivity (p=0.0117) and specificity (p<0.001) differed significantly. Conclusion A zero shot LLM achieved balanced performance, outperforming the structured phenotype on specificity while maintaining high sensitivity, supporting tiered ED screening for OUD.
Molina et al. (Fri,) studied this question.