Hydrological models are widely used to assess climate change impacts, but their performance can deteriorate under non-stationary climatic conditions. This study evaluates the robustness of a Long Short-Term Memory (LSTM) network versus two conceptual models (GR4J and SIMHYD) using the differential split-sample test within a large-sample experiment across 576 diverse Iranian catchments and a common objective function. All models exhibit the greatest performance decline when calibrated under wet conditions and validated under dry conditions. By explicitly separating climatic dissimilarity between calibration and validation periods from absolute validation-period aridity, the analyses show that LSTM performance is primarily controlled by differences in the aridity index, whereas conceptual model performance is more strongly governed by the absolute aridity of the validation climate. Additional analyses indicate that smaller, steeper catchments generally experience stronger performance deterioration, while larger, lower-slope catchments display more stable behaviour, particularly for the LSTM. Calibration-length experiments (4-16 years) show that increasing calibration duration substantially improves LSTM performance but yields only marginal gains for conceptual models, revealing a clear performance ceiling. Within this setup, approximately 12 years emerges as a practical threshold for robust comparison, although this value will vary with catchment properties, climate regime, data quality, and model architecture.
Jahanshahi et al. (Sun,) studied this question.