This paper presents an adaptive optimal consensus tracking control scheme for canonical nonlinear multi-agent systems (MASs) with unknown dynamics, employing an actor–critic reinforcement learning (RL) framework. The scheme integrates a sliding mode mechanism to suppress tracking errors and ensure consensus tracking between the followers and the leader. Additionally, optimal control is designed to find a Nash equilibrium in a graphical game. To address the intractability of obtaining an analytical solution for the coupled Hamilton–Jacobi–Bellman (HJB) equation, a policy iteration algorithm is utilized. Within this algorithm, a critic neural network (NN) approximates the gradient of the optimal value function, while an actor NN approximates the optimal control policy. Together, these networks form a compact actor–critic (AC) architecture that achieves optimal consensus tracking. Furthermore, the proposed method guarantees the boundedness of all closed-loop signals while ensuring consensus tracking. Finally, two simulations are conducted to verify the effectiveness and advantages of the proposed method.
Mo et al. (Tue,) studied this question.