Key points are not available for this paper at this time.
Multimodal remote sensing image classification has emerged as a key research area in remote sensing, with extensive applications in real-world scenarios. However, these images are collected by different sensors and contain multiple features such as spectrum, space, height and texture. Due to the differences in the characteristics of these data, existing methods have poor results in extracting and fusing heterogeneous features, which limits the improvement of classification performance. To address this problem, we propose a new heterogeneous feature extraction and fusion framework DTFNet, which utilizes the diffusion model and Transformer architecture. In the feature extraction stage, different networks are constructed to extract heterogeneous features while reducing redundancy. The dual-branch diffusion feature extraction (DBDFE) network based on the diffusion model is introduced to process data from different sensors, avoiding the limitation of extracting all features with a single network. In the feature fusion stage, the extracted diffusion features are fused with the original features to preserve the integrity of the original data. The cross-fusion transformer (CFT) module uses a convolutional neural network (CNN) to complete the local feature transformation and integration and models the long-range dependencies between heterogeneous features through cross-transformer encoders. Experimental results show that the classification accuracy of DTFNet on the three datasets reaches 92.38%, 80.08% and 95.02% respectively, which is significantly better than the existing state-of-the-art methods, demonstrating its effectiveness and superiority.
Ying et al. (Sat,) studied this question.