PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 9, 2026Remote Sensing0 citationsOpen Access

MPC-DETR: A Multi-Scale Patch Context Transformer for Small Object Detection in UAV Imagery

View Full Paper
QWQuanxiang WangZZZhaofa ZhouZZZhili Zhang

Key Points

  • This research aims to enhance the detection accuracy of small objects in UAV imagery, addressing challenges related to pixel size and complex backgrounds.
  • Proposed MPC-DETR framework integrating Local-Global Attention Fusion Module, Dilated Context-Aware Feature Interaction Module, and Patch-Guided Multi-scale Feature Fusion Module.
  • Implemented on VisDrone2019 dataset for performance evaluation against baseline models.
  • Included analysis on UAVDT and HIT-UAV datasets to assess generalization capability.
  • MPC-DETR achieved an mAP50 of 52.5% and mAP50–95 of 33.3%, surpassing baseline performance by 4.6 and 4.0 percentage points, respectively.
  • Enhanced feature representation through a lightweight multi-branch attention mechanism.
  • Demonstrated high generalization potential across different UAV-based small-object detection datasets.

Abstract

Small-object detection in unmanned aerial vehicle (UAV) imagery remains challenging because target objects often occupy only a few pixels, exhibit weak feature responses, and are easily obscured by complex backgrounds. These aspects significantly limit the effectiveness of end-to-end detection systems. To overcome these limitations and enhance the detection accuracy in challenging UAV settings, this paper proposes MPC-DETR, a Multi-scale Patch Context Transformer that is based on RT-DETR. To begin with, a Local-Global Attention Fusion Module (LGAF) is proposed to capture fine-grained local features and long-range semantic relations of small objects. LGAF enhances feature representation through a lightweight multi-branch synergistic attention mechanism while introducing limited computational overhead. Second, a Dilated Context-Aware Feature Interaction Module (DCFI) is proposed to enhance the discriminative capability of high-level features in cluttered backgrounds and densely populated small-object scenes. DCFI allows more efficient feature aggregation and contextual comprehension through multi-scale contextual modeling and scale-adaptive feature interaction. Third, a Patch-Guided Multi-scale Feature Fusion Module (PGMFF) is developed to create a patch-guided contextual fusion approach that combines shallow, high-resolution features with deeper semantic information. This process improves the maintenance and representation of fine object information and minimizes information loss in feature propagation. The experimental results on the VisDrone2019 dataset show that MPC-DETR has an mAP50 and mAP50–95 of 52.5% and 33.3%, respectively, which are 4.6 and 4.0 percentage points higher than the baseline model. Further analyses of the UAVDT and HIT-UAV datasets also support the high generalization potential of the suggested method to various UAV-based small-object detection problems. In general, the findings suggest that MPC-DETR provides precise, strong, and efficient small-object detection in complicated UAV images.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/6a782d6e2e1896536c840871https://doi.org/10.3390/rs18162650
Ask AI
Helpful
Bookmark
Share
View Full Paper