PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 2026Sensors0 citationsOpen Access

QA2FDet: Quality-Aware Adaptive Alignment Fusion Network for UAV RGBT Tiny Pedestrian Detection

View Full Paper
YTYifang TanLYLijun YuanCXChuanjiang Xie

Key Points

  • This research aims to enhance tiny pedestrian detection by addressing issues related to cross-modal misalignment and background interference in UAV imagery.
  • Developed QA2FDet, featuring spectrum-spatial decoupled enhancement, cross-modal correspondence mining, and prior-informed gated fusion.
  • Implemented discrete cosine transform for background disentanglement and deep semantic gating for noise suppression.
  • Utilized thermal-guided local asymmetric cross-attention for refining correspondences under slight spatial offsets.
  • QA2FDet achieved state-of-the-art performance on UAV RGBT detection benchmarks.
  • Demonstrated strong robustness in detecting tiny pedestrians even in challenging aerial scenes.

Abstract

Visible–thermal tiny pedestrian detection in UAV aerial images is crucial for online decision-making in urban security and disaster response. However, the extremely small scale and sparse distribution of pedestrians cause discriminative cues to be submerged by dominant low-frequency background and contextual redundancy during feature learning. Meanwhile, cross-modal spatial misalignment and spatially varying modality reliability hinder stable fine-grained correspondence, thereby degrading fusion quality. To address these issues, QA2FDet is proposed as a quality-aware adaptive alignment fusion network comprising three modules: spectrum-spatial decoupled enhancement module (SDE), cross-modal correspondence mining module (CCM), and prior-informed gated fusion (PGF). SDE leverages the discrete cosine transform to disentangle redundant low-frequency background information, while deep semantic gating propagates high signal-to-noise ratio details into shallow representations to enhance subtle cues of tiny pedestrians and suppress high-frequency noise. To establish fine-grained neighborhood correspondences under slight spatial offsets, thermal-guided local asymmetric cross-attention is designed in CCM. Finally, region-level quality and modality discrepancy are jointly modeled for adaptive cross-modal fusion in PGF. Extensive experiments on multiple UAV-based RGBT detection benchmarks demonstrate that QA2FDet achieves state-of-the-art performance and exhibits strong robustness in challenging aerial scenes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tan et al. (2026) studied this question.

synapsesocial.com/papers/6a1bd1745783ba022b6fcf91https://doi.org/10.3390/s26113443
Ask AI
Helpful
Bookmark
Share
View Full Paper