Dissolved organic carbon (DOC) is a critical water quality parameter (WQP) for water treatment processes (WTPs). This study trained and tested five machine learning (ML) models to predict DOC using the data collected from 113 WTPs in Ontario for 13 years (2008–2020). The Group-1 models had imputed data (zeroes), whereas the Group-2 models removed the empty rows. The stacking ensemble models (meta models) improved the prediction accuracy by combining ML models as classifiers. The Artificial Bee Colony and Gray Wolf Optimizer algorithms were applied for feature selection from 25 features. The Group-1 models had seven (color, nitrite, nitrate, sulfate, dissolved inorganic carbon, conductivity, chloride, and phosphate), and Group-2 models had four (color, nitrite, nitrate, sulfate, and alkalinity) WQPs. By incorporation of metaheuristic algorithms, features were optimized for predictive accuracy. In addition, combining the models through the stacking ensemble method and searching for optimal configurations improved model performance. In Group-1, performances of the main (R2 = 0.35–0.86) and meta (R2 = 0.87–0.97) models were lower than the Group-2 models (main: R2 = 0.9835–0.9973; meta: R2 = 0.9965–0.9997). Lower performances of Group-1 models indicated that data imputation needs careful attention. The models might be useful in dynamically adjustable coagulation processes.
Chowdhury et al. (Sat,) studied this question.