Background: Machine learning models for dental caries risk prediction require external validation to assess temporal generalizability. While feature stability is essential for clinical trust, few studies have tested whether behavioral predictors remain consistent across multiple population survey cycles. This study evaluated the temporal reproducibility of explainable machine learning models for dental caries risk across three National Health and Nutrition Examination Survey cycles spanning six years. Methods: This cross-sectional secondary analysis used National Health and Nutrition Examination Survey data from 2013–2014 (development set, n = 5,924), 2015–2016 (validation set one, n = 5,735), and 2017–2018 (validation set two, n = 5,533). Five supervised classification algorithms (logistic regression, random forest, XGBoost, support vector machine, and artificial neural network) were trained on 2013–2014 data and externally validated on subsequent cycles without retraining. Model performance was assessed using the area under the receiver operating characteristic curve. Feature importance stability was evaluated using Shapley Additive Explanations and Spearman rank correlation across cycles. Results: All models demonstrated declining discriminative performance over time. The best-performing algorithm in the development cycle (support vector machine, area under the receiver operating characteristic curve = 0.689) dropped to 0.614 in 2015–2016 and 0.622 in 2017–2018. XGBoost showed the strongest external validation performance (mean area under the receiver operating characteristic curve = 0.623). However, Shapley Additive Explanations feature rankings showed strong temporal stability (Spearman correlation = 0.949 for 2013–2014 versus 2015–2016; Spearman correlation = 0.847 for 2013–2014 versus 2017–2018). Age, dental visit frequency, and mouthwash use remained among the top predictors across all cycles, though the relative importance of socioeconomic and dietary variables declined in later cycles. Conclusions: Behavioral risk profiles for dental caries are temporally reproducible across six years, but their discriminative power weakens without model recalibration. These findings support the use of stable behavioral targets for population-level prevention while highlighting the need for periodic updating of artificial intelligence-driven risk screening tools to maintain predictive accuracy. Keywords: dental caries, machine learning, external validation, temporal stability, National Health and Nutrition Examination Survey, explainable artificial intelligence, Shapley Additive Explanations, behavioral risk factors, model degradation
Mohammad Mobassar Saleh Farhan Sikder (Sun,) studied this question.