Randomized trial demonstrates improved semantic segmentation in remote sensing, suggesting enhanced model adaptation.
High-resolution remote-sensing semantic segmentation requires models to simultaneously capture global scene semantics and preserve fine-grained local structures. Although satellite-pretrained vision foundation models provide strong transferable representations, the features extracted by a frozen backbone remain insufficiently adapted to dense prediction, particularly for representing high-frequency details and multiscale local patterns. In addition, correcting residual prediction errors with dense full-map refinement introduces substantial computational redundancy, since hard errors are typically concentrated in only a small subset of locations. To address these challenges, we propose ADVMSeg, an efficient remote-sensing semantic segmentation framework built upon a frozen satellite-pretrained DINOv3 backbone. Specifically, we introduce a Spatial-Frequency Adapter (SF-Adapter) to improve backbone-level dense feature adaptation by jointly modeling global frequency responses and multiscale local spatial details in a lightweight bottleneck space. We further design an Adaptive Sparse Refinement (ASR) module after the pixel decoder, which identifies hard regions from coarse predictions via uncertainty and boundary cues, and performs targeted local cross-attention refinement only on selected critical locations. Extensive experiments on GID-15, LoveDA, and ISPRS Potsdam validate the effectiveness of the proposed framework. Under the unified setting, ADVMSeg achieves 63.1% mIoU on GID-15, 63.5% mIoU on LoveDA, and 81.4% mIoU on ISPRS Potsdam. These results validate the effectiveness of jointly improving backbone-level feature adaptation and prediction-stage computation allocation under the evaluated setting of frozen DINOv3, and three representative remote-sensing semantic-segmentation datasets.
No takes yet. Share an insight, caveat, or question.
Ding et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: