ABSTRACT Forecasting Lake Tanganyika water level (WL) is crucial for flood and hydrological studies due to the linear, nonlinear, and irregular patterns. Nonlinearity is captured by deep learning but is challenged by data scarcity, noise interference, and hyperparameter tuning. To address these problems, this study proposes ARIMA-SSAb-LSTM and ARIMA-SSAa-LSTM to enhance forecast accuracy. While SSAb optimizes LSTM hyperparameters, SSAa reconstructs the data after removing noise and outliers. LSTM predicts nonlinear residuals produced by ARIMA when extracting the linear component. ARIMA-SSAb-LSTM model surpassed ARIMA-SSAa-LSTM, ARIMA-LSTM, SSAb-LSTM, WOA-LSTM, PSO-LSTM, LSTM, and ARIMA for both datasets for all metrics and achieved higher R2 than the others, with an increase on the train set of 5.06, 11.02, 13.17, 16.34, 19.98, 26.36, and 42.84%, respectively, and on the test set of 4.66, 12.51, 13.68, 25.19, 30.30, 46.52, and 62.40%, respectively. To assess generalization and learning, the models underwent training and testing. Clustering revealed the effect of seasonality on prediction performance, with ARIMA-SSAb-LSTM exhibiting strong performance across all hydrological regimes. These outcomes attest to the model's exceptional ability to predict reconstructed WL. The use of SSA in reconstruction and optimization increases prediction accuracy and facilitates efficient water resources management and flood protection in the vicinity of Lake Tanganyika.
Niyongabo et al. (2026) studied this question.