Dynamic programming with Q-learning based reinforcement learning optimization for multi unmanned aerial vehicles swarm path planning | Synapse