Chart review identified 34 patients with CAD and 103 with OSA, compared to only 16 and 33 identified by ICD codes, highlighting the utility of NLP for extracting clinical data beyond billing codes.
Observational (n=145)
No
NLP models show promise in extracting clinical information on cardiovascular disease and sleep apnea from EHRs more effectively than ICD codes alone, though challenges remain in context disambiguation.
Abstract Rationale Clinical notes in the electronic health record (EHR) contain important information beyond ICD codes which can be extracted using natural language processing (NLP). In an on-going World Trade Center (WTC)-related study on cardiovascular diseases (CVD) and obstructive sleep apnea (OSA), we first examined whether CVD and OSA-oriented information can be extracted from EHR by developing and testing the utility of an NLP-based machine learning model. Methods A physician manually reviewed clinical notes of 145 patients in the Mount Sinai Hospital EHR and annotated CVD events (coronary artery disease (CAD), myocardial infarction (MI), stroke, congestive heart failure (CHF)) and OSA status, extracted from sleep study results (in-lab and home-based). A subtype of NLP (bidirectional encoder representations from transformers) model was fine-tuned (n = 43, 80%/10%/10% train/validation/internal test set) with a named-entity recognition task, and tested on 102 subjects not seen by the model (external test set). Relevant ICD codes were obtained. Performance metrics were precision, recall, and F1 score. An F1 score greater than 80% was considered reliable. Results The cohort was male predominant (75.7%) with median age 53 years and median BMI 29.2 kg/m2. Chart review identified 34 patients with CAD, 6 with MI, 5 with CHF, and 4 with stroke. OSA was confirmed in 103/110 patients who had sleep studies done. We identified only 16 patients with CAD-relevant ICD codes and 33 patients with OSA-relevant codes. Due to low prevalence of non-CAD events in our cohort, we restricted classification to CAD and no-CAD. Performance metrics are reported in Table 1. False positives of the NLP model included inability to separate CAD diagnosis from patients’ medical history and family history and inability to distinguish CAD diagnosis from coronary calcium score. False negatives occurred when the model could not locate notes not based in our EHR (but available under Epic-Care Everywhere). Conclusion We highlight the utility and challenges of using NLP to extract accurate clinical CVD data to elucidate relationships with OSA. OSA has high prevalence in this WTC cohort. Clinical documentation for cardiac-related conditions was low as care was provided outside of our system, a common issue in a disintegrated healthcare network. Some patients with CAD or OSA diagnoses did not carry their respective ICD codes in their charts, underlining the importance of extracting clinical information beyond ICD codes. NLP shows promise in searching and returning accurate information for OSA. Additional data is needed to train the model. This abstract is funded by: CDC
Dong et al. (Fri,) conducted a observational in Cardiovascular diseases and obstructive sleep apnea (n=145). Natural language processing (NLP) machine learning model vs. ICD codes was evaluated on Precision, recall, and F1 score for extracting CAD and OSA status. Chart review identified 34 patients with CAD and 103 with OSA, compared to only 16 and 33 identified by ICD codes, highlighting the utility of NLP for extracting clinical data beyond billing codes.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: