Key points are not available for this paper at this time.
Accurate water quality prediction is essential for effective water pollution prevention and emergency responses. However, existing research on machine learning (ML)-based data assimilation methods remains limited, particularly in terms of addressing the combined impacts of climate change and anthropogenic activities. To address this gap, we proposed a novel ‘ML–Kalman filter (KF)’ data assimilation framework and evaluated its performance in the Dahei River Basin, a representative semi-arid watershed. Our results demonstrated significant improvements in predicting key water quality parameters, including total nitrogen (TN), total phosphorus (TP), and the permanganate index (COD Mn ), through the integration of KF with four ML models (LSTM, RF, XGBoost, and SVR). The accuracy enhancement ranged from 4.3 % to 17.6 %, with TP showing the most substantial improvement (9.2 %–17.6 %), followed by TN (6.4 %–11.1 %) and COD Mn (4.3 %–12.1 %). After assimilation, the models exhibited the following performance ranking for TN based on the coefficient of determination (R 2 ): LSTM–KF (R 2 = 0.909) > RF–KF (R 2 = 0.886) > SVR–KF (R 2 = 0.840) > XGBoost–KF (R 2 = 0.797), with similar trends observed for TP and COD Mn . The proposed framework demonstrates strong portability and applicability across different monitoring sections and temporal resolutions, offering a robust solution for regions with limited monitoring capabilities and challenging climatic conditions. These findings provide valuable data and technical support for advancing water pollution prediction and early warning systems, particularly for ecological and environmental departments operating in data-deficient regions. • The Kalman Filter effectively improved the prediction accuracy of machine learning models. • The prediction accuracy generally demonstrates the following results: LSTM-KF > RF-KF > XGBoost-KF > SVR-KF. • The LSTM-KF prediction method demonstrated considerable potential in future water environment monitoring and supervision. • This method exhibits good applicability and portability across datasets from various locations and time resolutions.
Gao et al. (Mon,) studied this question.