Transformer-based tracking methods have been widely studied in the field of visual object tracking. The long-range information capturing ability of the transformer improves the performance of the tracking network. However, the self-attention learning procedure in the transformer module neglects the local information, the target and the background around it, which can be beneficial for trackers to handle background clutter and deformation. In this paper, the local-global self-attention (LGSA) learning is proposed for the object tracking task, which obtains the local and global information simultaneously in one attention learning block. Based on the LGSA, the encoder and the decoder are designed to fuse the features corresponding to the template and search images. Additionally, two tracking networks, LGSAT-T and LGSAT-B instantiated with the proposed encoder and decoder are introduced. Exclusive experiments on the commonly used datasets, including OTB100, GOT-10K, LaSOT, and TrackingNet, demonstrate the effectiveness of LGSA, and indicate the state-of-the-art performance of the proposed tracking network. The code will be released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/lgao001/LGSAT</uri>.
No takes yet. Share an insight, caveat, or question.
Chen et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: