PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026IET conference proceedings.0 citations

PPO-Max obstacle avoidance optimization strategy of satellites

View Full Paper
XZXi ZhangGeneral CardiologyDNDan NiuSoutheast UniversityYCYang‐Yang ChenSoutheast University

Key Points

  • This research aims to optimize satellite trajectories for obstacle avoidance while minimizing energy consumption under specific constraints.
  • Constructed a constrained optimization model for obstacle avoidance in satellite operations.
  • Implemented a PPO-Max algorithm designed to reduce energy usage while navigating moving obstacles.
  • Simulated conditions using a three-body gravitational system and multi-scale reward functions.
  • PPO-Max achieved a 60% faster convergence rate during training.
  • Reduced energy consumption by approximately 25% compared to the standard PPO algorithm.

Abstract

This paper addresses the obstacle avoidance optimization problem of satellites, where satellites are constrained by fuel limitations, maximum velocity and minimum safety distance requirements, and obstacles are moving objects originating from another planet and transferred to the orbital region of the host planet. To simulate adversarial interference, obstacle behaviours are learned and fixed via opponent modelling, while their transfer trajectories are governed by a three-body gravitational system. Based on these fixed strategies and orbital dynamics, a constrained optimization model is constructed to solve for energy-optimal satellite avoidance trajectories under known obstacle interference. A so-called PPO-Max optimized strategy is designed to achieve minimizing the energy consumption of the satellite during obstacle avoidance. For the purpose of mitigating the effect of sparse reward functions on strategy learning, considering the optimization objective and constraints, multi-scale reward function is designed by combining survival reward, transfer reward, and fuel consumption. Simulations show that by combining the multi-scale reward function, the PPO-Max algorithm achieves a faster convergence rate of around 60% during training and reduces energy consumption by around 25% compared to the Proximal Policy Optimization (PPO) algorithm.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69fed0abb9154b0b82877b28https://doi.org/10.1049/icp.2026.1876
Ask AI
Helpful
Bookmark
Share
View Full Paper