Key points are not available for this paper at this time.
ABSTRACT The opioid overdose epidemic constitutes a critical public health crisis, necessitating advanced surveillance tools to enable timely intervention. Social media platforms provide a real‐time source of information on drug‐related behaviours. However, extracting structured knowledge from their informal, slang‐heavy and fragmented text presents significant technical challenges. While Named Entity Recognition (NER) enables the automated identification of drug‐related entities. Prior work has focused mainly on general biomedical domains, with limited exploration of domain‐specific scenarios. To address this gap, this study makes three key contributions. First, we present the first manually annotated, domain‐specific NER dataset for opioid overdose, comprising eight unique entity types, with a particular focus on routes of administration, sourced from Reddit posts spanning January 2021 to December 2023. Second, we provide a detailed description of the annotation process and guidelines, and systematically discuss challenges encountered during annotation, including slang, fragmented expressions and ambiguous language commonly found in social media posts. Third, we conduct an extensive series of experiments, including classical machine learning, deep learning with pretrained embeddings combined with CRF decoding, and transformer‐based models, evaluated at both token and entity levels. The proposed model achieved strong and balanced performance, with F1‐scores of 0.891 at both token and entity levels, highlighting the effectiveness of domain‐specific modelling for opioid overdose surveillance on social media.
Ahmad et al. (Tue,) studied this question.