The traditional waterway monitoring system has some problems such as delayed response, limited coverage and poor adaptability in a dynamic and complex environment. This paper proposes an intelligent path planning strategy for UAV waterway monitoring that integrates deep reinforcement learning (DRL) and meta-learning. First, construct a Partially Observable Markov Decision Process (POMDP) model to fuse multi-source sensor data and achieve dynamic environmental perception; Then, a compound reward function is designed, which takes into account the monitoring coverage, energy consumption, real-time performance and safety. The constraint reinforcement learning (CRL) mechanism is introduced, and the hard constraint of collision avoidance is explicitly embedded in the training process by Lagrange relaxation method, thus realizing Pareto optimal path generation. To enhance sample efficiency, a hierarchical experience replay mechanism is proposed. To rapidly adapt to unknown navigation routes, model-agnostic meta-learning (MAML) is employed to pre-train the strategy's initial parameters across diverse simulation environments. This enables the UAV to fine-tune for new routes using only a small amount of online data. The high fidelity experiment based on AirSim shows that the proposed method is significantly superior to A*, DDPG and basic SAC in terms of task coverage, collision times, energy consumption and completion time, and the collision rate is reduced by about 73%. The meta-learning mechanism can improve the coverage rate to over 90% in 1000 steps. The preliminary test of real UAV verifies the zero collision flight capability and engineering landing potential of the strategy.
Dongliang Cao (Sun,) studied this question.