In general, local visual-language (VL) trackers search targets around the previous bounding box by initial VL annotations. However, there is an inherent contradiction between the local searching perspective of the tracker and the orientation descriptions in language conducted under the global perspective. Furthermore, most methods only fuse modality information in a single stage, which tends to an insufficient relation modeling. To address these issues, we propose a Global Vision-Language Tracker (GVLTrack) with multi-stage modal fusion. First, it tracks the target in the entire image instead of local tracking based on previous results to resolve the above contradiction. Second, GVLTrack incorporates three modal interaction modules: consistent relationship modeling (CRM), VL-guided query initialization (VLQ), and recurrent cross-modal decoder (RC-Decoder) to fuse vision-language modality and refine the bounding box progressively comprehensively. We conduct extensive experiments on several benchmarks and achieve competitive performance, demonstrating the effectiveness of our approach. The code will be made publicly available as soon as it is accepted.
Tao et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: