Key points are not available for this paper at this time.
Vision transformers have recently been widely used in fine-grained visual classification. Most of the current studies utilizing vision transformers to mine key region features not only ignore the spatial connection between patches, but also additionally consider the background noise. Meanwhile, these methods learn the visual concepts of images individually, without considering the cue interactions between image pairs.To address these issues, this paper proposes a new spatial-aware feature enhancement network for fine-grained visual classification, which enables us to adaptively explore spatial contextual information in fine-grained salient regions, ignore interference from irrelevant background features, and discriminate subtle differences in similar targets by means of image pair interaction learning. Specifically, the proposed method includes two main modules: fine-grained spatial relationship and pairwise feature ensemble learning. The fine-grained spatial relationship module selects better discriminative regions by learning the mutual attention weights of different embedding layers, and subsequently adds location-dependent information to adaptively learn the neighborhoods of different regions through graph propagation. The pairwise feature ensemble learning module utilizes the mutual attention weights of the integrated image pairs to reduce confusion between fine-grained image pairs by guiding the feature interactions of the pairs from the perspective of each image individually through a gated residual mechanism. Finally, the complementary information from different transformer layers is added to the cross-layer feature boosting strategy for predicting classification results. We verify the effectiveness of this method on five widely used datasets and achieve excellent classification results. • SFE-Net explore spatial contextual information in fine-grained salient regions. • FSR selects better discriminative regions by learning the mutual attention weights. • PFEL guides feature interaction by image pair difference enhancement strategy. • Extensive experimental results demonstrate the effectiveness of the proposed method.
Kuang et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: