PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 24, 2026Sensors0 citationsOpen Access

MDCL-DETR: Multi-Domain Enhancement and Cross-Layer Feature Fusion for Small Object Detection

View Full Paper
THTianran HaoXZXi ZhangBZBing Zhou

Key Points

  • The aim is to enhance feature representation and contextual modeling for small object detection in UAV imagery.
  • Developed a multi-domain enhancement module for better feature distinction from background noise.
  • Utilized cross-layer feature extraction to maintain spatial details and integrate multi-scale features.
  • Implemented a gated Mamba fusion module for dynamic weighting of local details and global context.
  • Achieved mAP50 scores of 54.1% on VisDrone2019 dataset.
  • Achieved mAP50 scores of 56.2% on AI-TOD dataset.

Abstract

Small object detection in uncrewed aerial vehicle (UAV) imagery is hindered by limited pixels, insufficient detailed information, and strong background interference, leading to weak feature representation and poor contextual modeling. To address these issues, we propose a multi-domain enhancement and cross-layer feature fusion detection Transformer (MDCL-DETR) with progressive feature processing. First, a multi-domain enhancement module (MDEM) based on CSP (cross stage partial) structure is proposed, which fuses spatial and frequency-domain features in a lightweight manner to enhance object detail and global structures while effectively distinguishing object features from background interference. Second, a cross-layer feature extraction module (CLEM) is introduced to aggregate multi-scale features across layers, alleviate information loss caused by downsampling, and preserve spatial details of small objects while integrating high-level contextual semantics. Meanwhile, a gated Mamba fusion module (GMFM) is proposed, which adopts the Mamba architecture for long-range dependency modeling of multi-scale features and integrates a gating mechanism to realize the dynamic weighted fusion of local details and global context, further improving feature discriminability and global modeling capability. Finally, a fine-grained enhancement module (FGEM) is designed, which leverages feature reorganization and adaptive feature extraction to reinforce and compensate fine-grained features. Extensive experimental results validate the effectiveness and generalization of the proposed method, achieving mAP50 scores of 54.1% and 56.2% on the VisDrone2019 and AI-TOD datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hao et al. (2026) studied this question.

synapsesocial.com/papers/6a12969d48a0ea1665673844https://doi.org/10.3390/s26113305
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1HMF-DEIM: High-Fidelity Multi-Domain Fusion Transformer for UAV Small Object Detection2026 · 1 citations
  2. 2A Small Object Detection Transformer for UAV Remote Sensing Imagery via Multi-Scale Perception and Cross-Spatial-Frequency Domain Fusion2026 · 1 citations
  3. 3UMS-DET: A Frequency-Enhanced Transformer-Based UAV Multi-Scale Small Object Detector for Aerial Imagery2026
  4. 4ACD-DETR: Adaptive Cross-Scale Detection Transformer for Small Object Detection in UAV Imagery2025 · 10 citations
  5. 5MFE-DETR: Multimodal Feature-Enhanced Detection Transformer for RGB–Infrared Object Detection in Aerial Imagery2026 · 4 citations