This work proposes an interval incentive and predictive interpolation-based proximal policy optimization (IP-PPO) scheme for automated guided vehicle (AGV) path planning, to achieve fast convergence, strong generalization and sufficient smoothness in the control actions. Firstly, a predictive interpolation method is integrated into the traditional proximal policy optimization (PPO) framework. Then, an enhanced vector field histogram is constructed to generate the safe interval state, which is incorporated into the observation space. Furthermore, a safe interval-based reward function is formulated to enhance the obstacle avoidance performance. The interval incentive mechanism, which includes the safe interval state and interval reward function, is integrated into the predictive interpolation-based PPO to construct the IP-PPO framework. Finally, comparative simulations demonstrate that the proposed IP-PPO scheme exhibits superior learning efficiency, generalization performance, and strong robustness against model uncertainties while maintaining high smoothness of AGV path planning.
Liao et al. (Wed,) studied this question.