Accurate cloud and cloud-shadow segmentation is a crucial step in optical remote sensing image preprocessing, playing a significant role in subsequent applications such as land-cover classification and change detection. However, the complexity of cloud/shadow shapes and noise interference (e. g. , snow and ice, buildings, complex backgrounds, and atmospheric optics) make this task challenging. Although existing deep learning methods have achieved remarkable results in cloud segmentation tasks, a better balance between computational efficiency and segmentation accuracy is still needed. Traditional deep learning models have good detail and generalization capabilities due to their local feature extraction ability and spatial invariance, but they are relatively weak in processing global context information, leading to false positives and false negatives in complex scenarios. Encoders based on state space models (such as VMamba) can effectively capture global context through long-range dependency modeling, but there is still room for optimization in computational efficiency. Additionally, complex attention mechanisms (such as CBAM) can improve feature representation capability, but the large number of parameters limits the deployment efficiency of models. This paper conducts a systematic architectural exploration of the MCloudX cloud segmentation network, seeking a balance between efficiency and accuracy from three directions: backbone network modernization, encoder efficiency optimization, and attention mechanism lightweighting. Through comprehensive ablation experiments on SPARCS and L8-Biome datasets, we systematically evaluate the independent and synergistic effects of each component and validate them on Biome₃ and SPARCS datasets. Experimental results show that the proposed optimization configuration (ResNet50+LocalMamba+ECA-Net) significantly improves computational efficiency while maintaining comparable accuracy to the baseline. We name this optimization configuration LECloud, providing valuable empirical references for future research on efficient remote sensing segmentation architectures.
Lu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: