PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 25, 2026Sensors2 citationsOpen Access

A Systematic Study on Pretraining Strategies for Low-Label Remote Sensing Image Semantic Segmentation

View Full Paper
PLPeizhuo LiuHZHongbo ZhuXMXiaofei Mi

Key Points

  • The research aims to improve semantic segmentation performance for remote sensing images with limited labeled data through effective pretraining strategies.
  • Conducted systematic benchmarking of various self-supervised pretraining strategies.
  • Implemented a two-phase General-Purpose Pretraining (GPPT) followed by Domain-Adaptive Pretraining (DAPT).
  • Developed an Edge-Guided Masked Image Modeling (EGMIM) method for enhanced local feature learning.
  • The GPPT and DAPT framework significantly outperformed both single-phase pretraining methods and existing two-phase methods initialized from ImageNet.
  • Experiments on four RSI benchmarks showed consistent and substantial performance gains, especially in extreme low-label scenarios.
  • Provided mechanistic analyses explaining the synergistic effects of the pretraining phases.

Abstract

This paper addresses the critical challenge of semantic segmentation for remote sensing images (RSIs) under extremely limited labeled data. A high-quality initial model is paramount for downstream semi-supervised or weakly supervised learning paradigms, as it mitigates error propagation from the outset. We conducted a systematic investigation into self-supervised pretraining to serve this precise need. Within the low-label regime, we identify and tackle two pivotal factors limiting performance: (1) the domain shift between large-scale pretraining data and specific target tasks, and (2) the deficiency in local feature learning caused by large-window masking in visual foundation model (VFM) pretraining. To resolve these issues, we first benchmark various pretraining strategies, demonstrating that a two-phase General-Purpose Pretraining (GPPT) followed by Domain-Adaptive Pretraining (DAPT) framework is optimal, significantly outperforming both single-phase methods and the existing two-phase paradigm initialized from ImageNet. Subsequently, we propose an Edge-Guided Masked Image Modeling (EGMIM) method for the DAPT phase, which explicitly integrates edge priors to guide the masking and reconstruction process, thereby enhancing the model’s capability to capture fine-grained local structures. Extensive experiments on four RSI benchmarks validate the effectiveness of our approach, showing consistent and substantial gains, particularly in extreme low-label scenarios. Beyond empirical results, we provide in-depth mechanistic analyses to explain the synergistic roles of GPPT and DAPT.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/699e9106f5123be5ed04e528https://doi.org/10.3390/s26041385
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Task Specific Pretraining with Noisy Labels for Remote sensing Image Segmentation2024
  2. 2Leveraging Pretrained Priors for Weakly Supervised Semantic Segmentation of Remote Sensing Images2026
  3. 3Semi-supervised semantic labeling of remote sensing images with improved image-level selection retraining2024 · 3 citations
  4. 4A novel semi-supervised approach for semantic segmentation of aerial remote sensing images under limited ground-truth availability2024 · 1 citations
  5. 5CDEST: Class Distinguishability-Enhanced Self-Training Method for Adopting Pre-Trained Models to Downstream Remote Sensing Image Semantic Segmentation2024