Ensuring constraint satisfaction during the deployment of reinforcement learning (RL) controllers remains a key challenge for safety-critical systems. Model predictive shielding addresses this by verifying proposed actions through predictive models and replacing unsafe ones with a backup policy, but existing approaches can be overly conservative, computationally demanding, and difficult to design for nonlinear systems with uncertainty. We propose Adaptive Robust Model Predictive Shielding to overcome these limitations. First, we employ an approximate robust nonlinear model predictive controller as the backup policy, trained offline from multi-stage robust model predictive control data. This robust model predictive shielding approach retains safety under uncertainty while enabling real-time applicability. Second, we introduce an adaptive safety parameter in the RL observation space, allowing the agent to dynamically adjust its conservativeness. Our adaptive model predictive shielding method thus enhances safety and adapts to current uncertainty levels while avoiding excessive conservatism. When deployed with a safe backup policy, adaptive robust model predictive shielding retains safety under uncertainty and reduces unnecessary backup interventions. Simulation results for a nonlinear continuous stirred tank reactor with parametric uncertainty show that the proposed adaptive robust model predictive shielding approach reduces interventions of backup policies while still guaranteeing safety. This framework can be especially beneficial for safe RL of chemical processes where a combination of safety guarantees, high performance, and real-time feasibility is critical. • Propose Adaptive Robust Model Predictive Shielding for safe RL in process control. • Use offline-trained approximate robust NMPC backup for real-time safe deployment. • Embed adaptive safety parameter to reduce conservatism and improve performance.
No takes yet. Share an insight, caveat, or question.
Gerold et al. (2025) studied this question.