Type 2 diabetes mellitus (T2DM) involves progressive pancreatic beta cell dysfunction. Direct somatic-to-beta cell reprogramming, first demonstrated by Zhou et al. (2008) using the transcription factor trio PDX1-MAFA-NGN3 and subsequently extended by Furuyama et al. (2019) and Huang et al. (2023) to additional cell sources, represents an actively investigated autologous cell-replacement strategy. No published computational framework scores individual patients for reprogramming candidacy from routinely available clinical data. The closest methodological analogue we identified, Xu et al. (2025), predicts stem cell therapy outcome in plastic surgery from clinical variables using classical machine learning, but does not address diabetes, reprogramming biology, uncertainty quantification, or biological simulation. We present BetaForge-X, a machine learning pipeline integrating the Pima Indians Diabetes Dataset with twelve biologically informed synthetic gene expression features anchored to published GEO GSE15932 transcriptomic fold-changes (Marselli et al. , 2010). Six classifiers were trained with class-weight balancing to address diabetes class imbalance: Logistic Regression, Random Forest, Gradient Boosting, SVM, a deep neural network, and a feature-level self-attention scorer applied to this task for the first time to our knowledge, though the underlying attention mechanism itself follows established tabular-transformer architectures. Monte Carlo Dropout (n = 500 forward passes) provides per-patient uncertainty estimates, decomposed into aleatoric and epistemic components following Kendall and Gal (2017). Three SHAP explainer classes (Tree, Linear, Deep) provide model interpretability. Prior to finalising the manuscript, we conducted an eight-part pre-submission robustness audit: bootstrap confidence intervals (1000 resamples), Decision Curve Analysis (Vickers and Elkin, 2006), failure case characterisation, noise-injection robustness comparing clinical-only versus clinical-plus-gene feature sets, aleatoric/epistemic uncertainty decomposition, five-seed reproducibility testing, fifty-resample feature importance stability analysis, and age-tertile subgroup analysis. The ablation study demonstrates that synthetic gene features do not significantly improve classification AUC over clinical features alone (ΔAUC = −0. 0102, Wilcoxon p = 0. 968), clarifying that their contribution lies in constructing an interpretable reprogramming candidacy score and parameterising a patient-specific Hill-kinetics ODE simulation rather than in discriminative classification. Attention entropy analysis shows the FeatureAttentionScorer converges to 100. 0% of maximum uniform entropy at nₜrain = 537, consistent with the established literature on tree-based model dominance over from-scratch deep learning at this data scale (Grinsztajn et al. , 2022; Shwartz-Ziv and Armon, 2022) ; we identify prior-fitted architectures such as TabPFNv2 (Hollmann et al. , 2025) as the appropriate future direction rather than claiming the present architecture succeeds empirically. Subgroup analysis reveals that reprogramming candidacy scores are substantially age-confounded (mean score 0. 356 in the youngest tertile versus 0. 693 in the oldest), a finding disclosed explicitly rather than omitted. The best classifier (Logistic Regression, bootstrap 95% CI for AUC: 0. 715–0. 855) achieves clinically meaningful diabetic recall (0. 731) after class-weight balancing, compared to 0. 48 in an earlier unbalanced iteration of this pipeline. The framework relies on a single-cohort, female-only dataset and synthetic rather than measured transcriptomic features, has undergone no clinical or in-vitro validation, and exhibits the age confound noted above; these limitations are stated explicitly rather than deferred to a closing paragraph, and the work is positioned throughout as a computational hypothesis-generation tool situated at the clinical-prediction-to-computational-prioritisation stage of the translational pipeline, not as a validated clinical instrument. Its contribution lies in demonstrating that clinical machine learning, biologically anchored feature synthesis, uncertainty quantification, multi-class explainability, and mechanistic simulation can be combined into a single reproducible, openly documented pipeline for a therapeutic indication — beta cell reprogramming candidacy — for which, to our knowledge, no such integrated computational framework has previously been reported, together with a structured pre-submission robustness audit intended to make its scope and limitations explicit ahead of any future validation effort. Keywords: Type 2 Diabetes, Beta Cell Reprogramming, Machine Learning, Explainable AI, Uncertainty Quantification, Monte Carlo Dropout, SHAP, Hill-ODE Simulation, Decision Curve Analysis, Reproducibility
Suhail Ahmed Chandio (Fri,) studied this question.