Key points are not available for this paper at this time.
Multimodal remote sensing (RS) images exhibit distinct structure and distribution characteristics, making it challenging to design an effective multimodal RS image classification algorithm. Moreover, although existing deep learning-based methods have become the darling in the multimodal RS image classification, they usually lack effective exploration and explicit integration for generative information from different modalities. Aiming at the above challenges, a generative information-guided heterogeneous cross-fusion network with contrastive learning (GIHCN) is proposed for multimodal RS image classification. Firstly, to simulate the land-cover distributions from different modal data, a multimodal generative information learning architecture (MGILA) is constructed to capture the unsupervised heterogeneous distribution features. Secondly, to achieve bidirectional modeling between heterogeneous data and the reconstructed land-cover distributions, a heterogeneous data & generative information cross-attention module (HGCM) is designed to explore the complementarity between multimodal data and the reconstructed land-cover distributions. HGCM can provide the heterogeneous generative information for current modal data or provide the heterogeneous data support for current modal generative information, thereby obtaining cross-fusion sources with different attributes. Furthermore, we achieve the effective feature extraction for different cross-fusion sources by a designed multimodal contrastive learning framework (MCLF). Notably, to capture local information and long-range dependencies, a hybrid classification network with convolutional neural network and Mamba (CMNet) is proposed as the feature extraction backbone of each cross-fusion source to further improve the classification performance. Finally, we construct a joint multimodality loss function for MCLF, which can reduce the distribution difference between modalities while focusing on the information flow within and across the modality. Experimental results on four multimodal RS datasets confirm the effectiveness of GIHCN compared with other state-of-the-art methods. The source code will be released at https://github.com/ZJier/GIHCN.
Zhang et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: