Road extraction from high-resolution remote sensing images (HRSIs) provides critical support for downstream tasks such as autonomous driving path planning and urban planning. Although deep learning-based pixel-level segmentation methods have achieved significant progress, they still face challenges in handling occlusions caused by vegetation and shadows, and often exhibit limited model robustness and generalization capability. To address these limitations, this paper proposes the SAM2MS model, which leverages the powerful generalization capabilities of the foundational vision model, segment anything model 2 (SAM2). Firstly, an adapter-based fine-tuning strategy is employed to effectively transfer the capabilities of SAM2 to the HRSI road extraction task. Secondly, we subsequently designed a subtraction block (Sub) to process adjacent feature maps, effectively eliminating redundancy during the decoding phase. Multiple Subs are cascaded to form the multi-scale subtraction module (MSSM), which progressively refines local feature representations, thereby enhancing segmentation accuracy. During training, a training-free lossnet module is introduced, establishing a multi-level supervision strategy that encompasses background suppression, contour refinement, and region-of-interest consistency. Extensive experiments on three large-scale and challenging HRSI road datasets, including DeepGlobe, SpaceNet, and Massachusetts, demonstrate that SAM2MS significantly outperforms baseline methods across nearly all evaluation metrics. Furthermore, cross-dataset transfer experiments (from DeepGlobe to SpaceNet and Massachusetts) conducted without any additional training validate the model’s exceptional generalization capability and robustness.
Zhang et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: