Los puntos clave no están disponibles para este artículo en este momento.
Existing multi-modal UAV tracking methods typically rely on fixed-interval dynamic template update strategies to capture diverse target appearances, together with predefined thresholds to select high-quality search regions for template update. However, due to the irregular motion of targets and the complexity of real-world scenarios, such passive update mechanisms suffer from notable limitations. Fixed sampling intervals often fail to adequately capture appearance variations, while fixed threshold-based selection is insufficient to accommodate diverse imaging conditions, leading to ineffective updates or the introduction of noisy templates, thereby degrading tracking robustness and accuracy. To address these issues, we propose a search region-guided adaptive dynamic template update framework for robust multi-modal UAV tracking, aiming to improve both scene adaptability and target matching capability. Specifically, we design a Guided Template Selection Transformer, which dynamically matches templates conditioned on the current search region, enabling the tracker to autonomously select the most suitable template for the target’s current state. Furthermore, we introduce a Dynamic Threshold Module that adaptively adjusts template selection criteria according to different tracking scenarios, ensuring the reliability and contextual relevance of candidate templates. In addition, we develop a Dynamic Template Memory Module to maintain an ordered repository of target templates under different target states, providing a structured and high-quality template pool for the proposed selection mechanism. Extensive experiments on a standard multi-modal UAV tracking benchmark demonstrate that the proposed method significantly outperforms existing approaches, effectively overcoming the limitations of conventional fixed update strategies. Moreover, the proposed approach exhibits strong generalization capability across three additional multi-modal tracking datasets from typical surveillance scenarios.
Liu et al. (Tue,) studied this question.