Multimodal tracking is a crucial visual task that focuses on accurately locating specific targets in video frames. The primary challenge lies in effectively utilizing visual features to identify relevant positions. Existing methods often rely on advanced visual encoders and decoders to extract features from both visible and other modalities. However, due to the limited availability of multimodal data, relying solely on visual information is insufficient. Inspired by vision–language models, we propose the RGB-thermal (RGBT) tracking network with semantic generation and historical context (SHT). This approach addresses the lack of linguistic information in visual tracking and explores the semantic relationships between the target and its search area. Our approach utilizes large models to generate image descriptions, enhancing the target’s appearance information. Furthermore, it introduces the detail text visual focus (DTVF) module to improve the consistency between visual and textual data. In addition, we present a historical prompt generation method that combines historical foreground masks with visual features to provide precise cues for tracking purposes. The experimental results show that incorporating image descriptions and historical information significantly enhances multimodal tracking performance.
No takes yet. Share an insight, caveat, or question.
Gao et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: