Purpose This study utilizes interpretable machine learning to identify and prioritize key associated factors for adolescent obesity across individual, family, and school domains, as well as to establish specific risk thresholds that can inform targeted interventions. Methods Data were obtained from the China Education Panel Survey (CEPS), which included 7,397 adolescents. Six ML models (SVM, XGBoost, LightGBM, LR, RF, MLP) were developed and evaluated. The best-performing model was interpreted using SHAP analysis to assess feature contributions. Results The LightGBM model demonstrated the highest accuracy (0.8788). This study primarily focused on the accurate classification of adolescent obesity status within a clinical decision-making context. Consequently, accuracy was prioritized as the key metric for directly assessing the model’s overall classification performance. Key predictors of this model sedentary time, school ranking, academic workload, birth weight, body image, family economic status, school location, household registration, and physical activity. Among these, sedentary behavior emerged as the most significant predictor. Specific risk thresholds were identified, including sedentary time exceeding 5 h on weekends and birth weight greater than 4.0 kg. Conclusion This study underscores the utility of interpretable ML in identifying key predictors associated with adolescent obesity. The findings suggest that interventions might prioritize reducing sedentary behavior, the moderation of academic workload, and the enhancement of body image perception. Additionally, family and school environments play crucial roles in the prevention of obesity.
Huang et al. (Wed,) studied this question.