European beech forests are increasingly threatened by environmental stressors, leading to declines in vitality as well as tree mortality. Crown defoliation is widely used as an indicator of forest condition because it integrates the combined influences of site characteristics, climatic variability, and biotic pressures. Although machine learning models have been increasingly applied to predict defoliation patterns, model interpretability and the role of validation strategies in assessing predictive performance remain insufficiently explored. In this study, we compared traditional statistical models (generalized linear models, GLM; generalized additive models, GAM) with machine learning approaches, including Random Forest, XGBoost, and the Feature Tokenized Transformer (FT-Transformer), as well as ensemble models. Crown defoliation across the German federal states of Schleswig-Holstein, Lower Saxony, and Hesse was predicted using geo-ecological and climatic predictors. Model performance was evaluated under three validation schemes representing different generalization scenarios: random split, temporal hold-out (leave-one-year-out), and spatial hold-out (leave-one-location-out). Predictive accuracy was highest under random split validation (R² up to approximately 0.80) and decreased under temporal validation, while spatial validation produced substantially lower performance, indicating limited spatial transferability. Machine learning models consistently outperformed GLM and GAM across all validation schemes. SHapley Additive exPlanations (SHAP) indicated that crown defoliation patterns were primarily associated with site-related soil properties, particularly bulk density, while climatic predictors such as previous-year solar radiation contributed to interannual variability. These findings highlight the importance of validation design when evaluating predictive models and demonstrate both the potential and limitations of data-driven approaches for forest health monitoring.
Xu et al. (Sun,) studied this question.