Abstract Background Accurate risk prediction of severe coronary artery disease (CAD) is essential for early diagnosis and intervention. While clinical models are widely used, integrating machine learning (ML) and genetic data may enhance predictive performance. This study evaluates the effectiveness of ML models using clinical and genetic risk factors and validates findings in an independent external dataset (N=25,124). Purpose To assess the predictive performance of ML models trained on clinical and genetic risk factors for severe CAD, determine the added value of genetic data, and validate results in a large external dataset. Methods A discovery dataset of 983 severe CAD cases and 983 controls after downsampling was used to train six ML models: Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), Logistic Regression (LR), Linear Discriminant Analysis (LDA), and Decision Tree (DT). Features included clinical risk factors and 55 single nucleotide polymorphisms (SNPs) previously associated with CAD. An independent external dataset (N=25,124) was used for validation. Model performance was evaluated using AUC-ROC, PR-AUC, Accuracy, MCC, and Brier Score to assess classification and probability calibration. Results In the discovery dataset using clinical data alone, GBM achieved the highest recall (0.629) and AUC-ROC (73.26), making it the best classifier for severe CAD, while LDA demonstrated the highest accuracy (0.667) and precision (0.683), indicating better discrimination of non-CAD cases. SVM exhibited superior probability calibration with a Brier Score of 0.199, making it more reliable for individualized risk assessment. External validation confirmed these trends, with GBM achieving the highest AUC-ROC (88.99 88.19-89.78) and recall (0.822), while LDA maintained high accuracy (0.808) and precision (0.792). Notably, SVM had the highest recall (0.846) and demonstrated better calibration (Brier Score = 0.131), reinforcing its suitability for personalized risk estimation. When genetic data were incorporated, only marginal improvements were observed, with GBM’s recall increasing slightly to 0.641 and AUC-ROC to 73.63 in the discovery dataset. Feature importance analysis highlighted clinical risk factors, particularly age, BMI, LDL, and T2D, as the strongest predictors, reinforcing the dominance of clinical features in CAD classification. Conclusion(s) Clinical factors remain the strongest predictors of severe CAD, with genetic data providing minimal classification improvement. External validation confirms GBM as the best model for screening and SVM for individualized risk prediction. These findings highlight the importance of model selection based on application: GBM for population-wide screening and SVM for accurate risk estimation.
Hageh et al. (Sat,) studied this question.