Feature importance analysis is essential for interpreting machine learning models in diabetes mellitus (DM) risk prediction; however, existing interpretability methods often produce inconsistent feature rankings across models. This study proposes a unified ODE-inspired interpretability framework and an algorithmic decision procedure for robust feature selection by integrating contribution-based (SHAP), perturbation-based (permutation importance), and sensitivity-based feature importance measures. Multiple supervised machine learning models, including Logistic Regression, Random Forest, Gradient Boosting, Histogram Gradient Boosting, and Multilayer Perceptron, were trained and evaluated on a longitudinal biochemical and demographic dataset comprising 200 patients with three repeated visits (N = 600 observations). To preserve longitudinal integrity and avoid patient-level information leakage, grouped cross-validation was applied. A sensitivity-based feature importance formulation using finite-difference approximations enabled model-agnostic comparison across heterogeneous machine learning architectures. Stability, normalization, and cross-method agreement analyses were additionally introduced to evaluate consistency of feature rankings across models and interpretability methods. Experimental results consistently identified HbA1c as the dominant predictor, followed by lipid-related variables, age, and body mass index. Strong agreement was observed between ODE-inspired feature importance and SHAP analysis, whereas permutation importance demonstrated comparatively weaker agreement with sensitivity-based methods. The proposed framework further enabled systematic analysis of ranking stability, cross-method agreement, longitudinal sensitivity dynamics, and the introduction of an agreement-weighted Consensus Interpretability Score (CIS) for unified feature ranking across heterogeneous interpretability methods. The results demonstrate that integrating ODE-inspired sensitivity analysis with machine learning provides a robust, interpretable, and computationally scalable framework for feature importance assessment in diabetes risk prediction. The proposed approach offers a principled solution to inconsistent feature importance estimation and supports more reliable interpretation of biomedical machine learning models.
Knights et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: