PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 15, 2026Biomimetics2 citationsOpen Access

Multi-Head Attention Deep Q-Network with Prioritized Experience Replay for UAV Path Planning in Dynamic Environments: A Bio-Inspired Approach

View Full Paper
YLYan LiKunming University of Science and TechnologyXQXinjie QianYibin UniversityJZJiexin ZhangYibin University

Key Points

  • This research aims to enhance UAV path planning efficiency in dynamic environments using a novel AI architecture.
  • Developed a Multi-Head Attention Deep Q-Network integrated with Prioritized Experience Replay.
  • Used a 46-dimensional state space to represent environmental factors.
  • Implemented a curriculum learning strategy with 10 difficulty levels.
  • Conducted extensive simulations to evaluate performance.
  • Achieved a 96% task success rate for UAV target reaching without collisions or battery issues.
  • Improved convergence speed by 68% compared to baseline DQN.
  • Enhanced path efficiency by 17% and reduced energy consumption by 26%.

Abstract

Unmanned Aerial Vehicles (UAVs) have become widely used tools for different applications including surveillance, search and rescue, and package delivery. However, autonomous path planning in dynamic environments with moving obstacles, wind disturbances, and energy constraints remains a significant challenge. This paper proposes a novel Multi-Head Attention Deep Q-Network with Prioritized Experience Replay (MA-DQN + PER) that integrates bio-inspired attention mechanisms with deep reinforcement learning for efficient UAV path planning. Our approach features a 46-dimensional state space that captures all environmental information, including static obstacles, wind conditions, and energy status. The proposed Attention-QNetwork architecture uses four specialized attention heads to selectively focus on different aspects of the environment, including obstacle avoidance, target tracking and energy management, and wind compensation. To improve sample efficiency and convergence speed, we incorporate Prioritized Experience Replay (PER) as well as Prioritized Experience Replay (PER) with a sum-tree data structure to improve sample efficiency and convergence speed. A curriculum learning strategy that includes 10 difficulty levels is designed to progressively enhance the agent’s capabilities. Extensive simulations demonstrate that our MA-DQN + PER approach reaches a 96% task success rate (defined as the percentage of episodes where the UAV successfully reaches the target without collision or battery depletion), while the convergence speed was 68% quicker than that of the baseline DQN. Our method demonstrates superior performance in path efficiency (+17%), energy consumption reduction (−26%), and collision avoidance compared to state-of-the-art algorithms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69df2c01e4eeef8a2a6b0f32https://doi.org/10.3390/biomimetics11040268
Ask AI
Helpful
Bookmark
Share
View Full Paper