Short-term river water-level forecasting is essential for operational hydrology, supporting flood warning and water management. Although deep learning models such as Long Short-Term Memory (LSTM) networks have gained attention, classical statistical approaches including Autoregressive Integrated Moving Average (ARIMA) and Seasonal Autoregressive Integrated Moving Average (SARIMA) remain attractive due to their interpretability and efficiency. This study presents a controlled comparison between ARIMA/SARIMA and stacked LSTM models for 7-day-ahead water-depth forecasting using synthetic daily hydrographs representing normal, drought, and flood regimes. Model performance is assessed using a rolling-origin forecasting strategy that generates multiple overlapping predictions, reducing bias from short validation windows. Forecast skill is evaluated through standard error metrics and hydrology-oriented indicators, including the Global Forecast Skill Index (GFSI). Results show comparable median performance between SARIMA and LSTM across regimes, with no statistically significant differences detected by nonparametric tests. Apparent differences in flood conditions should be interpreted cautiously due to limited sample representation. Overall, increased model complexity does not inherently guarantee superior predictive skill in this univariate short-term setting, highlighting the importance of rigorous evaluation design in comparative forecasting studies.
Sîrbu et al. (2026) studied this question.