Multimodal remote sensing segmentation commonly integrates optical imagery with Digital Surface Models (DSMs) to improve land-cover understanding. However, existing methods often overlook a critical issue, namely the inconsistency between spatial-domain structural details and frequency-domain semantics, which can lead to misaligned features and degraded performance in complex scenes. To address this problem, we formulate spatial–frequency representation conflict as a unified multimodal learning problem and propose the Conflict-Aware Fusion Network (CAF-Net). Specifically, Cross-Modal Structure Guidance (CMSG) extracts DSM-derived high-pass structural cues and conditionally modulates optical features to improve boundary consistency. The Adaptive Cross-Frequency Module (ACFM) separates DCT coefficients using a fixed radial mask, adaptively reweights low- and high-frequency components, and performs cross-modal alignment at the highest encoder stage. Uncertainty-Aware Fusion (UAF) predicts pixel-wise relative reliability scores and normalizes them across modalities to suppress low-confidence responses. This coordinated design links spatial refinement, frequency alignment, and reliability-guided fusion instead of treating them as independent feature-enhancement operations. Experiments on the ISPRS Vaihingen and Potsdam datasets yield mIoU scores of 84.35% and 86.86%, respectively. The results indicate that explicitly modeling spatial–frequency discrepancies can improve multimodal segmentation accuracy and representation consistency.
Ma et al. (Wed,) studied this question.