Accurate maize yield prediction is essential for food security but remains challenging under small-sample and spatially heterogeneous conditions. To address this issue, a Partitioned Adaptive Data Synthesis (PADS) framework is proposed that integrates a partitioned version of the Synthetic Minority Over-sampling Technique for Regression with Gaussian Noise (SMOGN) and adaptive data selection to improve the quality of synthetic data for cross-regional yield prediction. Using 1,085 maize yield records from major maize-producing regions in China, PADS was evaluated against Global SMOGN under different augmentation levels. Results show that PADS produced synthetic data with improved distributional consistency and reduced discrepancy, leading to more stable gains in prediction for nonlinear models. Random Forest and Multilayer Perceptron achieved increases in R2 of approximately 0.03–0.04 and reductions in RMSE of 2–3%, while low-yield prediction improved more substantially, with tail RMSE and MAE reduced by over 13%. Further analyses indicate that these gains were mainly driven by improved synthetic data quality rather than by sample quantity alone, with moderate augmentation levels yielding the most stable benefits. These findings demonstrate that PADS offers an effective and transferable strategy for improving crop yield prediction under data-scarce and spatially heterogeneous conditions.
Xu et al. (Wed,) studied this question.