PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 17, 2025Drones4 citationsOpen Access

Off-Policy Deep Reinforcement Learning for Path Planning of Stratospheric Airship

View Full Paper
JXJiawen XieWHWanning HuangJMJinggang Miao

Key Points

  • RPL-TD3 enhances path planning speed by 62.5% compared to baseline methods.
  • Incorporating LSTM allows the model to consider historical states for improved decision-making.
  • A unique experience replay mechanism prioritizes high-value experiences to accelerate learning.
  • The method successfully meets kinematic and energy constraints in various simulations.

Abstract

The stratospheric airship is a vital platform in near-space applications, and achieving autonomous transfer has become a key research focus to meet the demands of diverse mission scenarios. The core challenge lies in planning feasible and efficient paths, which is difficult for traditional algorithms due to the time-varying environment and the highly coupled multi-system dynamics of the airship. This study proposes a deep reinforcement learning algorithm, termed reward-prioritized Long Short-Term Memory Twin Delayed Deep Deterministic Policy Gradient (RPL-TD3). The method incorporates an LSTM network to effectively capture the influence of historical states on current decision-making, thereby improving performance in tasks with strong temporal dependencies. Furthermore, to address the slow convergence commonly seen in off-policy methods, a reward-prioritized experience replay mechanism is introduced. This mechanism stores and replays experiences in the form of sequential data chains, labels them with sequence-level rewards, and prioritizes high-value experiences during training to accelerate convergence. Comparative experiments with other algorithms indicate that, under the same computational resources, RPL-TD3 improves convergence speed by 62.5% compared to the baseline algorithm without the reward-prioritized experience replay mechanism. In both simulation and generalization experiments, the proposed method is capable of planning feasible paths under kinematic and energy constraints. Compared with peer algorithms, it achieves the shortest flight time while maintaining a relatively high level of average residual energy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2025) studied this question.

synapsesocial.com/papers/68d45e4e31b076d99fa5e3fdhttps://doi.org/10.3390/drones9090650
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Multi-UAV Path Planning and Following Based on Multi-Agent Reinforcement Learning2024 · 55 citations
  2. 2Airships as Unmanned Platforms: Challenge and Chance2002 · 26 citations
  3. 3A novel DDPG method with prioritized experience replay2017 · 255 citations
  4. 4Experiential Systems Engineering Education Concept Using Stratospheric Balloon Missions2019 · 8 citations
  5. 5Virtual Sensoring of Motion Using Pontryagin’s Treatment of Hamiltonian Systems2021 · 46 citations