PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026IET conference proceedings.0 citations

PPO-Max obstacle avoidance optimization strategy of satellites

View Full Paper
XZXi ZhangDNDan NiuYCYang‐Yang Chen

Key Points

  • This research aims to optimize satellite trajectories for obstacle avoidance while minimizing energy consumption under specific constraints.
  • Constructed a constrained optimization model for obstacle avoidance in satellite operations.
  • Implemented a PPO-Max algorithm designed to reduce energy usage while navigating moving obstacles.
  • Simulated conditions using a three-body gravitational system and multi-scale reward functions.
  • PPO-Max achieved a 60% faster convergence rate during training.
  • Reduced energy consumption by approximately 25% compared to the standard PPO algorithm.

Abstract

This paper addresses the obstacle avoidance optimization problem of satellites, where satellites are constrained by fuel limitations, maximum velocity and minimum safety distance requirements, and obstacles are moving objects originating from another planet and transferred to the orbital region of the host planet. To simulate adversarial interference, obstacle behaviours are learned and fixed via opponent modelling, while their transfer trajectories are governed by a three-body gravitational system. Based on these fixed strategies and orbital dynamics, a constrained optimization model is constructed to solve for energy-optimal satellite avoidance trajectories under known obstacle interference. A so-called PPO-Max optimized strategy is designed to achieve minimizing the energy consumption of the satellite during obstacle avoidance. For the purpose of mitigating the effect of sparse reward functions on strategy learning, considering the optimization objective and constraints, multi-scale reward function is designed by combining survival reward, transfer reward, and fuel consumption. Simulations show that by combining the multi-scale reward function, the PPO-Max algorithm achieves a faster convergence rate of around 60% during training and reduces energy consumption by around 25% compared to the Proximal Policy Optimization (PPO) algorithm.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69fed0abb9154b0b82877b28https://doi.org/10.1049/icp.2026.1876
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A PPO-Based Air-Space Collaborative Monitoring Method for Maritime Search and Rescue2026
  2. 2Comparison of Optimization Methods for the Attitude Control of Satellites2024 · 1 citations
  3. 3Path Planning Optimization of Autonomous Robots Based on Deep Reinforcement Learning2026
  4. 4A POA-QPSO Hybrid Algorithm for Multi-Objective Optimization of Dual-Layer Walker Constellations2026 · 2 citations
  5. 5Long-Term Fuel-Optimal Collision Avoidance Maneuvers with Station-Keeping Constraints2024 · 14 citations