PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 8, 2026Chemical Engineering Science0 citationsOpen Access

Stable control and energy-saving optimization for the blast furnace process: A digital twin-guided reinforcement learning approach

View Full Paper
GZGuanwei ZhouMLMeng LiYYYaowei Yu

Key Points

  • The aim is to optimize the control of silicon content in hot metal for increased stability and efficiency in blast furnace operations.
  • Developed a hybrid digital twin combining a first-principles model with a dual-channel correction network.
  • Employed a Long Short-Term Memory-augmented Twin Delayed Deep Deterministic Policy Gradient agent for control optimization.
  • Used data from a commercial blast furnace for training and evaluating the agent's performance.
  • Achieved high predictive accuracy with the digital twin for silicon content.
  • Successfully stabilized silicon content within the target range of [0.3, 0.6] wt.%.
  • Identified a superior energy-saving strategy that reduced resource consumption compared to historical operations.

Abstract

The control of the hot metal silicon content is a central challenge for the stability and efficiency of blast furnace operations. Both conventional operator-driven control strategies and model-based methods, such as proportional–integral–derivative (PID) control and model predictive control (MPC), often struggle with the long-time delays, strong nonlinearities, and the difficulty of balancing short-term stability with long-term economic objectives in the blast furnace. To address these limitations, this paper proposes a novel control framework guided by a hybrid digital twin and optimized via deep reinforcement learning in an offline training setting. First, a hybrid digital twin was developed by coupling a first-principles mechanistic model (implemented in Aspen Plus) with a dual-channel residual correction network. This network is specifically designed to handle the multi-scale dynamics of the blast furnace, using a Gated Recurrent Unit (GRU) with an attention mechanism for slow-dynamic variables and a Multi-Layer Perceptron (MLP) pathway for fast-dynamic variables. Subsequently, a Long Short-Term Memory-augmented Twin Delayed Deep Deterministic Policy Gradient (LSTM-TD3) agent is trained within this environment. The agent’s recurrent architecture explicitly addresses the long delays of the process, while the TD3 algorithm ensures stable and effective policy optimization. Experimental results, based on data from a commercial blast furnace, demonstrate that the digital twin achieves high predictive accuracy. The trained LSTM-TD3 agent successfully stabilizes the hot metal silicon content within the target range of 0.3, 0.6 wt.%. Furthermore, it discovers a superior energy-saving strategy, reducing key resource consumption compared to historical operations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2026) studied this question.

synapsesocial.com/papers/6a265bb6ad53cfb9357c5315https://doi.org/10.1016/j.ces.2026.124366
Ask AI
Helpful
Bookmark
Share
View Full Paper