The dominant paradigm in machine learning-based disease risk prediction assesses models exclusively through cross-sectional accuracy metrics, leaving unanswered whether temporal model predictions respond to simulated lifestyle intervention in a physiologically plausible manner. This paper proposes behaviouralrobustness(BR) as a new evaluation criterion fortemporal disease risk models, complementing accuracy by quantifying trajectory-level plausibility under controlled intervention. Our core contribution is the BR metric itself—not a novel model architecture. We instantiate BR using a hybrid LSTM–Transformer framework and evaluate it across two synthetic cohorts used as controlled metric-validation conditions (Pima, N=768; NHANEScalibrated, N=2,000) and two real clinical datasets where BR is computed on actual patient trajectories (UCI Diabetes CGM,N=70; Diabetes 130-US Hospitals EHR, N=5,246). Against five published baselines, the non-personalised Transformer achieves the highest AUC on UCI (0.897) and NHANES (0.695). Statistically significant behavioural inertia is observed on all four datasets (Wilcoxon p<10−8); BR is consistently low (0.003–0.022), within the range implied by published intervention trial effect sizes. We make no claim that synthetic BR values substitute for real-world evidence; the real-data results on Diabetes-130 (N=5,246, p=1.71×10−284) constitute the primary empirical contribution. External validation on MIMIC-III and UK Biobank is identified as the essential next step
nanda et al. (Mon,) studied this question.