Rapid and accurate mapping of earthquake-triggered landslides from satellite imagery is critical for emergency response and hazard assessment, yet remains challenging due to irregular boundaries, extreme size variations, and atmospheric noise. This paper proposes BiFusion-LDSeg, a novel bi-directional fusion enhanced latent diffusion framework that synergistically combines CNN-Transformer architectures with generative diffusion models for robust landslide segmentation. The framework introduces three key innovations: (1) a dual-encoder with Bi-directional Attention Gates (Bi-AG) enabling sophisticated cross-modal feature calibration between local CNN textures and global Transformer context; (2) a conditional latent diffusion process operating in learned low-dimensional landslide shape manifolds, reducing computational complexity by 100× while enabling inference with only 10 sampling steps versus 1000+ in standard diffusion models; and (3) a boundary-aware progressive decoder employing multi-scale reverse attention mechanisms for precise boundary delineation. Comprehensive experiments on three earthquake datasets from Sichuan Province, China (Lushan Mw 7.0, Jiuzhaigou Mw 6.5, Luding Mw 6.8) demonstrate superior performance, outperforming state-of-the-art methods by 7–13% in IoU and 5–7% in DSC across all three datasets. The framework exhibits exceptional noise robustness, strong cross-dataset generalization, and inherent uncertainty quantification, enabling reliable deployment for post-earthquake landslide inventory mapping at regional scales.
Shi et al. (2026) studied this question.