Remote sensing images, which obtain surface information from aerial or satellite platforms, are of great significance in fields such as environmental monitoring, urban planning, agricultural management, and disaster response. However, due to the complex and diverse types of ground coverage and significant differences in spectral characteristics in remote sensing images, achieving high-quality semantic segmentation still faces many challenges, such as blurred target boundaries and difficulty in recognizing small-scale objects. To address these issues, this study proposes a novel deep learning model, MKF-NET. The fusion of KAN convolution and Vision Transformer (ViT), combined with the multi-scale feature extraction and dense connection mechanism, significantly improves the semantic segmentation performance of remote sensing images. Experiments were conducted on the LoveDA dataset to systematically evaluate the segmentation performance of MKF-NET and several existing traditional deep learning models (U-net, Unet++, Deeplabv3+, Transunet, and U-KAN). Experimental results show that MKF-NET performs best in many indicators: it achieved a pixel precision of 78.53%, a pixel accuracy of 79.19%, an average class accuracy of 76.50%, and an average intersection-over-union ratio of 64.31%; it provides efficient technical support for remote sensing image analysis.
Ye et al. (Fri,) studied this question.