Rapid and accurate landslide detection from remotely sensed data is fundamental but significant work for large-scale geological disaster prevention and reduction. Vision transformers (ViT) have emerged as dominant solutions alongside CNN in landslide modeling, but they face challenges in terms of global and local information interactions, while the homogeneity between landslide areas and surrounding environments is difficult to distinguish solely from spectral images. This paper proposes a hierarchical perception-guided deep learning framework to detect landslides with multi-source remotely sensed data. Following the encoder‒decoder structure, we introduce a hybrid encoding backbone with the wavelet transform to model the spectral and terrain features. It is formulated with dual-branch configurations, where the CNN-based local perception module has lightweight network blocks, and the ViT-based global one comprises efficient multi-stage fusion blocks. We then explore developing the fusion mechanism with a cross-attention paradigm to promote interactive learning between heterogeneous global and local landslide representations. Furthermore, a U-shaped multi-scale feature decoding module with a comprehensive objective function is proposed to produce high-quality landslide segmentations. Extensive experiments on two benchmark landslide detection datasets demonstrate its competitive performance by (98.70%, 78.37%, 73.48%) and (99.68%, 80.55%, 76.10%) in terms of Acc, F1-score, and mIoU, respectively.
Zhong et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: