Key result
The HCM-VAr-Risk machine learning model effectively identified hypertrophic cardiomyopathy patients with ventricular arrhythmias, achieving a C-index of 0.83.
Why the study?
Existing clinical risk stratification models for SCD in HC utilize only a few variables, prompting assessment of whether data-driven machine learning considering a wider range of variables can effectively identify HC patients with ventricular arrhythmias.
Does a machine learning model using electronic health records effectively identify hypertrophic cardiomyopathy patients with ventricular arrhythmias?
Observational (n=711)
Does a machine learning model using electronic health records effectively identify hypertrophic cardiomyopathy patients with ventricular arrhythmias?
Effect estimate: C-index 0.83
A machine learning model incorporating 22 clinical variables from electronic health records demonstrated strong predictive performance (C-index 0.83) for identifying ventricular arrhythmias in patients with hypertrophic cardiomyopathy, outperforming traditional risk models.
May support EHR-based ML for VA risk stratification in HCM; hypothesis-generating, requires prospective validation before adoption.
Clinical risk stratification for sudden cardiac death (SCD) in hypertrophic cardiomyopathy (HC) employs rules derived from American College of Cardiology Foundation/American Heart Association (ACCF/AHA) guidelines or the HCM Risk-SCD model (C-index ∼0.69), which utilize a few clinical variables. We assessed whether data-driven machine learning methods that consider a wider range of variables can effectively identify HC patients with ventricular arrhythmias (VAr) that lead to SCD. We scanned the electronic health records of 711 HC patients for sustained ventricular tachycardia or ventricular fibrillation. Patients with ventricular tachycardia or ventricular fibrillation (n = 61) were tagged as VAr cases and the remaining (n = 650) as non-VAr. The 2-sample ttest and information gain criterion were used to identify the most informative clinical variables that distinguish VAr from non-VAr; patient records were reduced to include only these variables. Data imbalance stemming from low number of VAr cases was addressed by applying a combination of over- and undersampling strategies. We trained and tested multiple classifiers under this sampling approach, showing effective classification. We evaluated 93 clinical variables, of which 22 proved predictive of VAr. The ensemble of logistic regression and naïve Bayes classifiers, trained based on these 22 variables and corrected for data imbalance, was most effective in separating VAr from non-VAr cases (sensitivity = 0.73, specificity = 0.76, C-index = 0.83). Our method (HCM-VAr-Risk Model) identified 12 new predictors of VAr, in addition to 10 established SCD predictors. In conclusion, this is the first application of machine learning for identifying HC patients with VAr, using clinical attributes. Our model demonstrates good performance (C-index) compared with currently employed SCD prediction algorithms, while addressing imbalance inherent in clinical data.
No takes yet. Share an insight, caveat, or question.
Bhattacharya et al. (2019) conducted an observational in Hypertrophic cardiomyopathy (n=711). HCM-VAr-Risk Model (machine learning) vs. ACCF/AHA guidelines or HCM Risk-SCD model was evaluated on Separation of ventricular arrhythmia (VAr) from non-VAr cases (C-index 0.83). The HCM-VAr-Risk machine learning model effectively identified hypertrophic cardiomyopathy patients with ventricular arrhythmias, achieving a C-index of 0.83.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: