Key points are not available for this paper at this time.
This article investigates critical challenges in optimal output tracking control, including residual tracking errors, potential instability caused by discount factors, and premature convergence due to inefficient termination criteria. We tackle these issues by developing a data-driven parallel Q-learning algorithm. Specifically, a utility function directly linked to system states is proposed to avoid the instability and error amplification issues in traditional discounted approaches. In addition, the algorithm uses dual lightweight controllers that use convergence properties to enhance learning efficiency. Based on dual controllers, a novel termination criterion is introduced to prevent premature convergence during the training process. Numerical simulations demonstrate that the proposed method eliminates tracking errors, accelerates convergence compared with traditional algorithms, and ensures stable convergence across diverse system dynamics.
Wang et al. (Tue,) studied this question.