ABSTRACT The integration of hyperspectral imagery and LiDAR data offers promising potential for multimodal feature learning in urban land cover classification. However, effectively extracting and fusing these heterogeneous data sources to fully leverage their complementary strengths remains a challenging problem. To address this issue, we propose the Cascade Encoder–Decoder Fusion Network (CEDFNet), a novel multimodal framework designed for high‐precision urban land cover classification. CEDFNet employs two parallel Cascade Encoder–Decoder Networks (CEDNets) as its backbone, where the cascaded architecture enables progressive multi‐scale feature integration and improves the discrimination of land cover patterns across different spatial resolutions. In addition, the model incorporates two specialized modules: the Complementary Feature Focusing Module (CFFM), which enhances cross‐modal complementarity and produces high‐quality fused representations, and the Dense Attention Branch (DAB), which adaptively captures both low‐level and high‐level attentive cues to further strengthen feature expressiveness. Experimental evaluations on the Houston 2018 and MUUFL Gulfport datasets demonstrate that CEDFNet consistently outperforms state‐of‐the‐art baseline models, confirming its effectiveness and robustness in complex urban environments with diverse land cover distributions.
Wang et al. (Thu,) studied this question.