PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Sensors0 citationsOpen Access

Reinforcement Learning-Based Cloud-Aware HAPS Trajectory Optimization in Soft-Switching Hybrid FSO/RF Cooperative Transmission System

View Full Paper
BCBeibei CuiSCShanyong CaiLWLiqian Wang

Key Points

  • This research aims to optimize HAPS trajectories in hybrid FSO/RF systems by integrating deep reinforcement learning for cloud avoidance.
  • Developed a joint optimization framework integrating soft-switching between FSO and RF links.
  • Utilized deep reinforcement learning (DRL) with a proximal policy optimization (PPO)-based agent.
  • Employed rateless codes (RCs) for simultaneous transmission and developed a reward-shaped DRL training process.
  • Carried out simulations using realistic ERA5 data to evaluate performance.
  • The RC-PPO agent achieved higher throughput compared to the HS-PPO baseline.
  • Trajectory smoothness improved significantly with RC-PPO.
  • The soft-switching approach minimized link transitions and reduced instability in the transmission.

Abstract

Space–air–ground systems employing free-space optical (FSO) communication leverage high-altitude platform stations (HAPS) to deliver seamless and ubiquitous connectivity. Although FSO links offer high capacity, they are highly susceptible to cloud extinction, which severely degrades link availability. Hybrid FSO/radio-frequency (RF) transmission and cloud-aware HAPS trajectory optimization can enhance resilience. However, the conventional cloud-aware hybrid FSO/RF transmission system based on hard-switching (HS) between the FSO and RF links leads to frequent link transitions and unstable throughput. To address these challenges, we propose a joint optimization framework that integrates soft-switch between FSO and RF links with deep reinforcement learning (DRL) for HAPS trajectory optimization. Soft-switching based on rateless codes (RCs) enables simultaneous transmission over both links, where the receiver accumulates packets until successful decoding with a single feedback. The feedback frequency of RC is sparse, which avoids feedback storms but also poses challenges to HAPS trajectory optimization. The DRL agent proactively optimizes HAPS trajectories to avoid cloud cover and maintain link availability. To address the sparse feedback of RCs for DRL training, a reward-shaped proximal policy optimization (PPO)-based agent is developed to jointly optimize throughput and trajectory smoothness. Simulations using realistic ERA5 data show that RC-PPO achieves higher throughput and smoother trajectories compared to the HS-PPO baseline.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cui et al. (2026) studied this question.

synapsesocial.com/papers/6984345ff1d9ada3c1fb2689https://doi.org/10.3390/s26030948
Ask AI
Helpful
Bookmark
Share
View Full Paper