Diabetes is due to a variety of interacting factors including behavioral, metabolic and environmental ones. The large number of biomedical data collected in an increasingly multi-modal manner has given rise to many potential methods in bioinformatics to unify disparate types of data and provide insights into biological processes. In this paper we propose a bioinformatics-based approach to predict diabetes risk by unifying clinical biomarkers with digital lifestyle factors, social economic status, and environmental exposure. Our goal was to develop a machine learning-based pipeline on a curated data set containing 1879 individuals, where all steps were clearly defined (data integration, data transformation), and the importance of each input variable could be assessed (SHAP values). We identified the ensemble-based models (Random Forest and XGBoost) to have the best predictive capability for the task at hand; however, we also found that the results provided insight into the biological mechanisms involved in the onset of diabetes and demonstrated the potential utility of using bioinformatics approaches to support the design of algorithms capable of providing interpretable recommendations for early intervention and tailored health care strategies.
Harbaoui et al. (Fri,) studied this question.