Diabetes mellitus is a chronic condition that frequently leads to severe complications that are difficult to detect in their early stages using conventional clinical monitoring. This paper presents a data-driven framework for predicting multiple diabetes-related complications using structured electronic health record data while ensuring clinically meaningful explainability. The proposed approach adapts a pretrained electronic health record foundation model to operate on static patient data and integrates it with classical machine learning baselines to address class imbalance, feature sparsity, and interpretability challenges. A multi-label prediction setting covering eight common diabetes complications is evaluated using a real-world dataset from a regional diabetes center in the United Arab Emirates. Synthetic data generation and clinical constraint enforcement are applied to improve robustness for underrepresented outcomes, while feature selection is guided by model importance and attribution-based explanations. The best-performing configuration, a weighted ensemble combining a low-rank adapted Hyena-based foundation model with a tree-based predictor, achieved an average F1-score of 0.77, an average recall of 0.85, and an example-based F1-score of 0.71, outperforming all individual models. In addition, this ensemble produced the most stable explanations under input perturbations, indicating improved consistency of dominant clinical risk drivers. These results demonstrate that explainable foundation model-based ensembles can deliver accurate, robust, and clinically transparent risk prediction for diabetes complications.
Joseph et al. (Sat,) studied this question.