Spacecraft safe proximity, as a critical component of on-orbit servicing missions, primarily encounters the following two challenges: the partial observability of the environment surrounding the service spacecraft and the necessity to evade uncertain obstacles. A safe reinforcement learning algorithm based on a graph neural network is proposed to address the constrained Markov decision problem in partially observable scenarios for spacecraft safe proximity missions. A graph neural network mechanism is introduced to solve the problem of dynamic variations in the quantity and location of obstacles in the observation area of the service spacecraft. The graph attention network is used to facilitate the extraction of feature information from the graph structure, which is then utilized as input for the subsequent reinforcement learning algorithm. The Soft Actor–Critic–Lagrangian algorithm is adopted to deal with the problems of tuning reward function parameters and balancing safety and optimality. By introducing Lagrange multipliers, the constrained optimization problem is transformed into an unconstrained optimization problem. In order to verify the effectiveness of the algorithm proposed in this paper, a spacecraft safe proximity environment model with dynamic obstacles is constructed, and the GAT-SACL algorithm proposed in this paper is validated by the Monte Carlo shooting method. The results show that the GAT-SACL algorithm possess excellent exploratory characteristics and delivers significant advantages in balancing optimality and safety.
Zhou et al. (Thu,) studied this question.