PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 24, 2017PLoS ONE309 citationsOpen Access

Predicting diabetes mellitus using SMOTE and ensemble machine learning approach: The Henry Ford ExercIse Testing (FIT) project

MAManal AlghamdiMAMouaz H. Al‐MallahSKSteven J. Keteyian

Key Points

  • To evaluate the performance of machine learning algorithms and class-imbalance techniques in predicting incident diabetes from cardiorespiratory fitness records.
  • Analyzed data from N=32,555 patients without baseline coronary artery disease or heart failure who underwent treadmill stress testing (1991–2009) with 5-year follow-up.
  • Selected 13 key predictors from 62 candidate variables using clinical evaluation, multiple linear regression, and information gain ranking.
  • Applied Synthetic Minority Oversampling Technique (SMOTE) to address outcome imbalance and built a voting ensemble of Naïve Bayes Tree, Random Forest, and Logistic Model Tree classifiers.
  • A total of 5,099 patients (15.7%) developed incident diabetes over the 5-year follow-up period.
  • The voting ensemble classifier trained with SMOTE achieved a high predictive performance with an AUC of 0.92.

Abstract

Machine learning is becoming a popular and important approach in the field of medical research. In this study, we investigate the relative performance of various machine learning methods such as Decision Tree, Naïve Bayes, Logistic Regression, Logistic Model Tree and Random Forests for predicting incident diabetes using medical records of cardiorespiratory fitness. In addition, we apply different techniques to uncover potential predictors of diabetes. This FIT project study used data of 32,555 patients who are free of any known coronary artery disease or heart failure who underwent clinician-referred exercise treadmill stress testing at Henry Ford Health Systems between 1991 and 2009 and had a complete 5-year follow-up. At the completion of the fifth year, 5,099 of those patients have developed diabetes. The dataset contained 62 attributes classified into four categories: demographic characteristics, disease history, medication use history, and stress test vital signs. We developed an Ensembling-based predictive model using 13 attributes that were selected based on their clinical importance, Multiple Linear Regression, and Information Gain Ranking methods. The negative effect of the imbalance class of the constructed model was handled by Synthetic Minority Oversampling Technique (SMOTE). The overall performance of the predictive model classifier was improved by the Ensemble machine learning approach using the Vote method with three Decision Trees (Naïve Bayes Tree, Random Forest, and Logistic Model Tree) and achieved high accuracy of prediction (AUC = 0.92). The study shows the potential of ensembling and SMOTE approaches for predicting incident diabetes using cardiorespiratory fitness data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alghamdi et al. (2017) studied this question.

synapsesocial.com/papers/69dd5f497808b00a4799d4achttps://doi.org/10.1371/journal.pone.0179805
Ask AI
Helpful
Bookmark
Share
View Full Paper