Randomized trial compares precipitation imputation methods in Türkiye, suggesting hybrid models improve accuracy.
Precipitation records from meteorological stations frequently contain gaps caused by sensor, power, or transmission failures, creating uncertainty in hydrological, agricultural, and water-resources applications. This study compared two simple baselines (station-specific monthly climatological mean and temporal linear interpolation), deterministic spatial interpolation, direct reanalysis-based replacement, and machine-learning methods for daily precipitation imputation. Daily precipitation from 14 stations in Eastern and Southeastern Türkiye during 1985–2014 was evaluated using an independent final-test set formed by stratified random masking of 15% of complete observations; the remaining 85% was used for calibration, SHapley Additive exPlanations (SHAP) analysis, cross-validation, and hyperparameter optimization. ERA5-Land variables were transferred to the stations, precipitation was calibrated by Empirical Quantile Mapping, and leakage-controlled Kriging estimates were incorporated as predictors in XGBoost, LightGBM, Random Forest, Support Vector Regression, and Multilayer Perceptron models. The station-month climatological mean (RMSE = 5.4820 mm; NSE = 0.0527) and temporal linear interpolation (RMSE = 5.7059 mm; NSE = −0.0262) performed substantially worse than optimized Kriging and IDW. The full-hybrid LightGBM model achieved the best performance (RMSE = 3.2001 mm; MAE = 0.9814 mm; Pearson r = 0.8317; NSE = 0.6772), whereas direct ERA5-EQM replacement was less accurate (RMSE = 5.2252 mm; NSE = 0.1394). Combining local observations, spatial information, and ERA5-Land covariates therefore improved daily precipitation imputation in the study region.
No takes yet. Share an insight, caveat, or question.
TEKTAŞ et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: