This paper addresses the obstacle avoidance optimization problem of satellites, where satellites are constrained by fuel limitations, maximum velocity and minimum safety distance requirements, and obstacles are moving objects originating from another planet and transferred to the orbital region of the host planet. To simulate adversarial interference, obstacle behaviours are learned and fixed via opponent modelling, while their transfer trajectories are governed by a three-body gravitational system. Based on these fixed strategies and orbital dynamics, a constrained optimization model is constructed to solve for energy-optimal satellite avoidance trajectories under known obstacle interference. A so-called PPO-Max optimized strategy is designed to achieve minimizing the energy consumption of the satellite during obstacle avoidance. For the purpose of mitigating the effect of sparse reward functions on strategy learning, considering the optimization objective and constraints, multi-scale reward function is designed by combining survival reward, transfer reward, and fuel consumption. Simulations show that by combining the multi-scale reward function, the PPO-Max algorithm achieves a faster convergence rate of around 60% during training and reduces energy consumption by around 25% compared to the Proximal Policy Optimization (PPO) algorithm.
Zhang et al. (2026) studied this question.