Machine learning-based bleeding risk models demonstrated statistically superior discrimination compared to established scores in AF and VTE, though effect sizes were modest (ΔAUC 0.05–0.15).
Systematic Review (n=464,523)
Do machine learning-based bleeding risk models improve the prediction of bleeding events compared to conventional clinical scores in anticoagulated patients with AF and VTE?
Machine learning models offer modestly superior discrimination for predicting bleeding risk in anticoagulated AF and VTE patients compared to traditional clinical scores, though their routine clinical utility remains to be established.
Effect estimate: ΔAUC 0.05-0.15
Background: Accurate prediction of bleeding events in patients receiving oral anticoagulants remains a key challenge in the management of atrial fibrillation (AF) and venous thromboembolism (VTE). Machine learning (ML) algorithms have emerged as powerful tools that capture complex, nonlinear interactions among risk factors, potentially offering superior accuracy. Objectives: To synthesize evidence comparing ML-based bleeding risk models with conventional clinical scores in anticoagulated AF and VTE populations. Methods: We conducted a systematic review with narrative synthesis of studies published between 2015 and 2025 applying ML algorithms to predict bleeding events in anticoagulated AF or VTE patients. Results: Thirteen studies were identified (seven AF and six VTE), including 464,523 participants in total. ML algorithms such as random forest (RF), extreme gradient boosting (XGBoost), and neural networks consistently outperformed traditional tools. In AF, AUCs ranged from 0.64 to 0.76 compared to 0.52–0.61 for HAS-BLED. In VTE, ML models achieved 0.59–0.91 versus 0.61–0.65 for RIETE or VTE-BLEED. Deep learning ensembles reached the highest AUCs (>0.8). Conclusions: ML-based bleeding risk models demonstrated statistically superior discrimination compared to established scores in both AF and VTE contexts, but effect sizes were modest (ΔAUC 0.05–0.15) and clinical utility remains uncertain. Broader validation, calibration assessment, and demonstration of impact on clinical outcomes are necessary before routine adoption.
Teo et al. (Fri,) conducted a systematic review in Atrial fibrillation and venous thromboembolism (n=464,523). Machine learning-based bleeding risk models vs. Conventional clinical scores (HAS-BLED, RIETE, VTE-BLEED) was evaluated on Bleeding events prediction (AUC) (ΔAUC 0.05-0.15). Machine learning-based bleeding risk models demonstrated statistically superior discrimination compared to established scores in AF and VTE, though effect sizes were modest (ΔAUC 0.05–0.15).