PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 16, 2025Remote Sensing28 citationsOpen Access

STMSF: Swin Transformer with Multi-Scale Fusion for Remote Sensing Scene Classification

YDYingtao DuanCSChao SongYZYifan Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

Emerging vision transformers (ViTs) are more powerful in modeling long-range dependences of features than conventional deep convolution neural networks (CNNs). Thus, they outperform CNNs in several computer vision tasks. However, existing ViTs fail to encounter the multi-scale characteristics of ground objects with various spatial sizes when they are applied to remote sensing (RS) scene images. Therefore, in this paper, a Swin transformer with multi-scale fusion (STMSF) is proposed to alleviate such an issue. Specifically, a multi-scale feature fusion module is proposed, so that features of ground objects at different scales in the RS scene can be well considered by merging multi-scale features. Moreover, a spatial attention pyramid network (SAPN) is designed to enhance the context of coarse features extracted with the transformer and further improve the network’s representation ability of multi-scale features. Experimental results over three benchmark RS scene datasets demonstrate that the proposed network obviously outperforms several state-of-the-art CNN-based and transformer-based approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Duan et al. (2025) studied this question.

synapsesocial.com/papers/6a7d3bad149bd8e2c57497f3https://doi.org/10.3390/rs17040668
Ask AI
Helpful
Bookmark
Share
View Full Paper