Key points are not available for this paper at this time.
Reinforcement learning (RL) enables flying ad-hoc networks (FANETs) to choose the next hop unmanned aerial vehicles (UAV s) with shorter routing path, but may raise the retransmission rate and fails to guarantee the quality of service (QoS) under the high mobility and fast fading channels. In this paper, we propose an RL based routing scheme that optimizes both the routing and the power allocation to protect the latency QoS and save routing energy consumption of the FANET. Based on the routing history, the channel conditions, the battery level and the shared knowledge from the neighbors, this scheme formulates the routing policy distribution with safe exploration to select the stable path and thus reduce the retransmission rate. Specifically, the risk value with respect to end-to-end latency constraint is designed to evaluate the routing policy and reduce the exploration probability of the high-latency routing. Based on the distributed value function approach, the learning parameter such as the state value functions shared among neighbors is exploited to accelerate the routing process and enhance the routing stability under the dynamic network topology. Simulation results verify the routing performance gain of our proposed scheme over the benchmark.
Qi et al. (Mon,) studied this question.