PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 11, 2025PLoS ONE5 citationsOpen Access

Reinforcement learning for UAV flight controls: Evaluating continuous space reinforcement learning algorithms for fixed-wing UAVs

View Full Paper
HKHasan Raza KhanzadaAMAdnan MaqsoodABAbdul Basit

Key Points

  • Reinforcement learning algorithms outperformed classical PID controllers in flight stability and responsiveness, especially during disturbances.
  • The Soft Actor-Critic algorithm converged in 400 episodes, achieving steady-state error under 3% to demonstrate its effectiveness.
  • Five continuous-space reinforcement learning algorithms were compared for their applicability to UAV flight dynamics under uncertain conditions.
  • The analysis provides insights for selecting suitable reinforcement learning algorithms for modern UAV control systems.

Abstract

Flight controls are experiencing a major shift with the integration of reinforcement learning (RL). Recent studies have demonstrated the potential of RL to deliver robust and precise control across diverse applications, including the flight control of fixed-wing unmanned aerial vehicles (UAVs). However, a critical gap persists in the rigorous evaluation and comparative analysis of leading continuous-space RL algorithms. This paper aims to provide a comparative analysis of RL-driven flight control systems for fixed-wing UAVs in dynamic and uncertain environments. Five prominent RL algorithms that include Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO) and Soft Actor-Critic (SAC) are evaluated to determine their suitability for complex UAV flight dynamics, while highlighting their relative strengths and limitations. All the RL agents are trained in a same high fidelity simulation environment to control pitch, roll and heading of the UAV under varying flight conditions. The results demonstrate that RL algorithms outperformed the classical PID controllers in terms of stability, responsiveness and robustness, especially during environmental disturbances such as wind gusts. The comparative analysis reveals that the SAC algorithm achieves convergence in 400 episodes and maintains a steady-state error below 3%, offering the best trade-off among the evaluated RL algorithms. This analysis aims to provide valuable insight for the selection of suitable RL algorithm and their practical integration into modern UAV control systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khanzada et al. (2025) studied this question.

synapsesocial.com/papers/68e9b1d9ba7d64b6fc132d36https://doi.org/10.1371/journal.pone.0334219
Ask AI
Helpful
Bookmark
Share
View Full Paper