Relative humidity (RH) is an important meteorological factor that affects both the climate system and human activities. However, the existing observational station data are insufficient to meet the requirements of regional scale research. Machine learning methods offer new avenues for high precision RH estimation, but the performance of different algorithms in complex geographical environments still needs to be thoroughly evaluated. Based on Chinese observational station data from 2011 to 2020, this study systematically evaluated the performance of three methods for estimating RH: the generalized linear mixed model (GLMM), random forest (RF) and the XGBoost algorithm. The results of ten-fold cross validation indicate that the two machine learning methods are significantly superior to the traditional GLMM. Among them, RF performed the best (the determinant coefficient (R2) = 0.73, root mean square error (RMSE) = 8.85%), followed by XGBoost (R2 = 0.72, RMSE = 9.07%), while the GLMM performed relatively poorly (R2 = 0.58, RMSE = 11.08%). The model performance shows significant spatial heterogeneity. All models exhibit high correlation but relatively large errors in the northern regions, while demonstrating low errors yet low correlation in the southern regions. Meanwhile, the model performance also shows significant seasonal variations, with the highest accuracy observed in the summer (June to September). Among all features, dew point temperature (Td) aridity index (AI) and day of year (DOY) are the main contributing factors for RH estimation. This study confirms that the RF model provides the highest accuracy in RH estimation.
Yao et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: