Background/Objectives: Type 2 diabetes mellitus (T2DM) is a multisystemic disease with overlapping metabolic, renal, and cardiovascular effects. Within the Diabetic@ project, which aims to characterize individuals with T2DM using real-world data extracted from electronic health records (EHRs), this substudy sought to develop a predictive model for two-year heart failure (HF) risk. Methods: Multicenter, retrospective study including T2DM individuals across eight Spanish hospitals (2013–2018). Data were extracted exclusively from EHRs’ unstructured free text using clinical natural language processing (cNLP) and mapped to SNOMED CT. At inclusion, individuals were categorized as having or not prevalent HF (pHF). Predictive modeling was performed in non-pHF to assess two-year risk of developing HF, termed incident HF (iHF). Logistic regression (LR), decision trees, random forest, and XGBoost were compared, selecting for accuracy and interpretability. Results: Of 588,756 individuals with T2DM, 84,197 (14.3%) had pHF. Among non-pHF, 353,371 (60%) were used for model development (90.7% training, 9.3% validation). iHF occurred in 13.6% of the training set and 11.4% of the validation set. Ischemic heart disease was present in 16.2% overall, 37.9% in pHF, and 12.6% in non-pHF. Glycosylated hemoglobin data was rarely reported (<15%). LR achieved the best performance (AUC-ROC 0.73) using 27 predictors. Reduced 12- and clinically refined 9-predictor models performed similarly, with the latter implemented in a web-based tool. Conclusions: Unstructured data from EHRs enabled development of a two-year HF risk model for individuals with T2DM, underscoring the potential of cNLP for risk stratification across the cardiovascular–renal–metabolic spectrum.
Navarro-González et al. (Sat,) studied this question.