Key points are not available for this paper at this time.
Evaluating supervised models for estimating reference evapotranspiration (ET o ) is crucial when there is a domain shift, meaning the models should be tested at stations not used during training. In such cases, an external k-fold validation can be performed, reserving different stations for testing in each iteration. However, this simple approach may include training data from stations that differ substantially from the testing station, which can reduce the model's generalizability due to domain shift. Therefore, this study focused on transfer-learning (TL) based modeling of ET o to enhance the performance of temperature-based gene expression programming (GEP) models for spatial (external) modeling. Specifically, transductive TL, which emphasizes domain selection via an instance-based approach, was employed by applying preliminary clustering methods for the first time in the literature. Various initial clusters of training stations were defined based on factors such as the aridity index (IA), the continentality index (CICU), principal component analysis (PCA), the self-organizing maps (SOM), and k-means clustering. K-fold validation was conducted within each cluster, which provided intra-cluster and cross-cluster assessments. The models achieved lower accuracy when a global k-fold validation was used without preliminary clustering for source domain selection. The findings revealed the importance of the similarity between source (training) and target stations in model performance, indicating that preliminary clustering may be beneficial before data splitting in supervised machine learning development. When station clustering is appropriate, the intra-cluster approach generally yields more accurate results than the cross-cluster approach.
Kazemi et al. (Sat,) studied this question.