Based on the process characteristics and production data of converter steelmaking, this study proposes an endpoint phosphorus content prediction model using K-means clustering and Bayesian-optimized Stacking ensemble learning. Firstly, the K-means clustering method was applied to analyze converter smelting data, grouping samples with similar process characteristics into homogeneous sub-clusters. Then, within each sub-cluster, a Stacking ensemble model was established with Random Forest (RF), Support Vector Machine (SVM), Extreme Gradient Boosting (XGB), Gradient Boosting Machine (GBM), and Light Gradient Boosting Machine (LGBM) as base learners, and linear regression as the meta-learner. To validate the model's performance, the comparative experiments were conducted between a single Stacking model and individual base models after K-means clustering, with all models optimized via Bayesian hyperparameter tuning. Experimental results demonstrate that the Stacking model enhanced by K-means clustering and Bayesian optimization achieves the highest prediction accuracy. When prediction errors were constrained within ±0.005% and ±0.003%, the hit rates for endpoint phosphorus content reached 94.84% and 84.06%, respectively. This model provides accurate prediction of converter endpoint phosphorus content, offering valuable technical guidance for phosphorus control in actual production.
Liu et al. (Thu,) studied this question.