Key points are not available for this paper at this time.
With the rapid development in semantic segmentation in remote sensing (RS) imagery, the balance between segmentation accuracy and computational efficiency has been recognized as a key challenge. Most existing approaches have predominantly been based on spatial features, while the potential of frequency-domain information has remained underexplored. To address this limitation, a Spatial-Frequency Feature Fusion Network (SF3Net) is proposed, aiming to achieve accurate segmentation with reduced computational cost. The framework is composed of two core modules: the Frequency Feature Stereo Learning (FFSL) module, designed to extract frequency features from multiple orientations, and the Spatial Feature Aggregation Module (SFAM), developed to enhance spatial feature representation. In addition, a Feature Selection Module (FSM) is incorporated to retain shallow features from the encoder, thereby compensating for detail loss during downsampling. Comprehensive experiments demonstrate the effectiveness of SF3Net: it achieves mean Intersection-over-Union (mIoU) scores of 80.040% and 83.934% on the ISPRS-Potsdam and ISPRS-Vaihingen datasets, respectively, and 88.657% on a self-constructed farmland dataset. These results consistently surpass those of state-of-the-art spatial–frequency fusion methods. Overall, SF3Net provides an efficient and effective paradigm for jointly modeling spatial and frequency-domain features in remote sensing semantic segmentation.
He et al. (Mon,) studied this question.