Key points are not available for this paper at this time.
This paper addresses the challenge of multisource feature fusion for hyperspectral images (HSI) and light detection and ranging (LiDAR)/synthetic aperture radar (SAR) in complex scene classification tasks. A multimodal cross-scale transformer network (MMCTNet) is proposed to tackle this issue. The model employs a spatial self-attention (SSA) module to enhance intra-modal spatial dependency modeling, while a multiscale adaptive fusion (MSAF) module achieves cross-modal semantic complementation. Furthermore, a transformer encoder combined with a cross-attention mechanism facilitates global semantic interaction and feature collaboration. Experiments conducted on four public datasets-MUUFL, Augsburg, Berlin, and 2018Houston-demonstrate that MMCTNet achieves superior performance over existing methods in terms of overall accuracy (OA), average accuracy (AA), and the Kappa coefficient.
Gong et al. (Mon,) studied this question.