Drone-based RGBT person detection facilitates critical applications such as search and rescue, owing to its high maneuverability and inherent capability to mitigate visual occlusion. However, despite the complementary nature of RGBT systems, existing detectors often overlook the specific impact of occlusion during the fusion process, leading to feature contamination and subsequent detection failures. In this work, we address this limitation by formally defining two categories of occlusion: “soft occlusion,” where targets remain partially visible in at least one modality, and “hard occlusion,” which involves complete obstruction. To tackle these challenges, we propose Unveiling Occluded Targets (UOT), a novel multi-modal fusion framework that implements a Quality–Occlusion Arbitration (QOA) mechanism. By leveraging both quality-related and occlusion-related cues, UOT dynamically arbitrates the fusion process to maximize information recovery from the clearer modality. Extensive experiments on the RGBTDronePerson and VTUAV-det datasets demonstrate significant improvements, achieving an mAP50all of 53.42% and an mAP50tiny of 54.70% in densely occluded scenes. Qualitative analysis further confirms UOT’s robustness in reliably identifying targets obstructed by sparse foliage.
Gui et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: