Cyber-physical systems (CPS) play a pivotal role in industrial automation, transportation, and critical infrastructure, where meeting stringent timing constraints is essential to ensure operational safety and efficiency. While reinforcement learning (RL) has shown promise in synthesizing controllers for time-critical applications, existing approaches often prioritize speed (as soon as possible, or ASAP) without explicitly addressing deadline compliance. This misalignment can lead to unsafe or suboptimal behaviors, which are unacceptable in industrial contexts requiring both safety and reliability under hard deadlines. For example, with inappropriate rewards, a control policy for a industrial robot can be encouraged to be less safe but fast instead of being steady and meeting the deadlines. To address this challenge, we investigate the relationship between ASAP behavior and deadline-safe behavior, introducing a novel Markov decision process formulation (R-MDP) that includes time-awareness while preserving the Markov property. We propose a reward design method that systematically encourages deadline compliance and guarantees safety in reach-avoid tasks. Our approach is validated on multiple benchmarks, including linear and nonlinear systems representative of industrial applications, such as DC motor control and real-time attitude control. Experimental results demonstrate the efficacy of our method in achieving deadline-safe control while maintaining system safety, offering a reliable solution for industrial CPS where failure to meet deadlines can have severe consequences.
Liu et al. (Wed,) studied this question.