A TabNet deep-learning pipeline integrating proteomics, body composition, and EHR significantly outperformed conventional machine learning models in predicting type 2 diabetes progression.
Observational
Does a multimodal deep-learning pipeline (TabNet) improve T2DM risk prediction compared to conventional machine learning models in a Middle Eastern cohort?
A multimodal deep-learning framework integrating proteomics and body composition outperforms conventional machine learning in predicting T2DM risk, highlighting demographic-specific drivers.
Longitudinal analysis of high-dimensional biomedical data offers a transformative window into the transition from health to metabolic disease. Despite advancements in predictive modeling, the integration of temporal proteomic shifts with clinical phenotypes for type 2 diabetes mellitus (T2DM) risk assessment remains underdeveloped, particularly within underrepresented Middle Eastern populations. We developed an interpretable, deep-learning pipeline utilizing tabular network (TabNet), an attention-based architecture designed for tabular data, to analyze a longitudinal cohort from the Qatar biobank (QBB). The framework integrates multimodal inputs, including proteomic signatures, dual-energy X-ray absorptiometry (DXA) body composition metrics, and electronic health records (EHR). Model transparency was ensured through built-in feature selection and SHAP (SHapley Additive exPlanations) values. The performance of the TabNet model was rigorously benchmarked against gradient-boosting machines, support vector machines, and random forest. Our findings demonstrate that the TabNet significantly outperforms conventional machine learning benchmarks in capturing the non-linear trajectories of diabetes progression. The model identified a distinct cluster of proteomic markers and visceral adiposity metrics that serve as early indicators of T2DM. Age-and-gender-stratified analyses revealed significant demographic divergence in risk profiles; specifically, the predictive weight of inflammatory markers and visceral fat mass exhibited profound variance across age cohorts, suggesting the need for life-stage-specific screening protocols. This study establishes an explainable artificial intelligence (XAI)-driven framework that bridges the gap between molecular discovery and clinical risk stratification. By uncovering the demographic-specific drivers of T2DM, our results provide a foundation for personalized preventative medicine and highlight the efficacy of attention-based deep learning in managing complex, longitudinal biobank data.
Khan et al. (Mon,) conducted a observational in Type 2 diabetes mellitus (T2DM). TabNet deep-learning pipeline (proteomics, DXA, EHR) vs. Conventional machine learning models (gradient-boosting machines, support vector machines, random forest) was evaluated on Prediction of type 2 diabetes progression. A TabNet deep-learning pipeline integrating proteomics, body composition, and EHR significantly outperformed conventional machine learning models in predicting type 2 diabetes progression.