ABSTRACT This paper introduces a model‐free non‐zero‐sum game framework for safe and optimal vehicle platooning control, incorporating both control and state constraints. The cost function for each vehicle is enhanced with a control barrier function (CBF) and input constraint functions. This enhancement ensures safety and optimality, effectively transforming the constrained non‐zero‐sum game into an unconstrained problem. We propose an off‐policy integral reinforcement learning (IRL) algorithm, underpinned by the policy iteration (PI) method, to achieve Nash equilibrium and maintain system optimality. To facilitate the approximation of control inputs and cost functions, an actor‐critic network is employed, which utilizes data collected from the system. The effectiveness of this approach is demonstrated through simulations, validating the proposed method's capability to achieve safe optimal vehicle platooning control effectively.
Liu et al. (Fri,) studied this question.