PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 21, 20250 citationsOpen Access

UAV-CodeAgents: Scalable UAV Mission Planning via Multi-Agent ReAct and Vision-Language Reasoning

View Full Paper
OSOleg SautenkovSkolkovo Institute of Science and TechnologyYYYasheerah YaqootSkolkovo Institute of Science and TechnologyMMMuhammad Ahsan MustafaSkolkovo Institute of Science and Technology

Key Points

  • This work aims to develop a multi-agent framework for autonomous UAV mission generation using vision-language models.
  • Developed UAV-CodeAgents leveraging ReAct paradigm for mission planning
  • Utilized satellite imagery and natural language inputs for trajectory generation
  • Implemented a reactive thinking loop for adaptability
  • Fine-tuned model on 9,000 annotated satellite images to enhance spatial grounding
  • Achieved average mission creation time of 96.96 seconds
  • Demonstrated success rate of 93% in mission execution
  • Lower decoding temperature (0.5) resulted in higher planning reliability

Abstract

We present UAV-CodeAgents, a scalable multi-agent framework for autonomous UAV mission generation, built on large language and vision-language models (LLMs/VLMs). The system leverages the ReAct (Reason + Act) paradigm to interpret satellite imagery, ground high-level natural language instructions, and collaboratively generate UAV trajectories with minimal human supervision. A core component is a vision-grounded, pixel-pointing mechanism that enables precise localization of semantic targets on aerial maps. To support real-time adaptability, we introduce a reactive thinking loop, allowing agents to iteratively reflect on observations, revise mission goals, and coordinate dynamically in evolving environments. UAV-CodeAgents is evaluated on large-scale mission scenarios involving industrial and environmental fire detection. Our results show that a lower decoding temperature (0.5) yields higher planning reliability and reduced execution time, with an average mission creation time of 96.96 seconds and a success rate of 93%. We further fine-tune Qwen2.5VL-7B on 9,000 annotated satellite images, achieving strong spatial grounding across diverse visual categories. To foster reproducibility and future research, we will release the full codebase and a novel benchmark dataset for vision-language-based UAV planning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sautenkov et al. (2025) studied this question.

synapsesocial.com/papers/69473b64db9c958d0dfca7bchttps://doi.org/10.48550/arxiv.2505.07236
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs2026
  2. 2General-Purpose Aerial Intelligent Agents Empowered by Large Language Models2025 · 1 citations
  3. 3MiniUAV-VLA: A Compact Vision–Language–Action Model for Cooperative Multi-UAV Search and Elimination via MARL Expert Distillation2026
  4. 4UAV-VLPA*: A Vision-Language-Path-Action System for Optimal Route Generation on a Large Scales2025
  5. 5Multimodal AI for UAV: Vision–Language Models in Human– Machine Collaboration2025 · 5 citations