Aerial scene recognition has progressed substantially with deep learning methods for RGB and hyperspectral imagery; however, existing approaches typically operate on single modalities or rely on explicit multimodal fusion, limiting scalability, flexibility, and deployment in heterogeneous sensing environments. To address this limitation, we propose a sensor-agnostic semantic representation learning framework that formulates multimodal learning as the unification of semantic representations rather than feature-level fusion. The proposed architecture employs modality-specific encoders and projection heads to map spatial and spectral–spatial features into a shared semantic embedding space, enabling modality-invariant representation learning while preserving discriminative characteristics of each sensing modality. A composite objective integrating cross-spectral alignment, intra-class compactness regularization, and prototype-based semantic anchoring is introduced to enforce consistent embedding geometry and improve class separability across modalities. A unified classifier operating within this shared space enables reliable inference from a single modality input without requiring paired data or explicit fusion. Extensive evaluations on multiple benchmark datasets, including Houston 2013 for cross-modality RGB–hyperspectral analysis, UC Merced for independent RGB aerial scene classification, and Indian Pines for hyperspectral land-cover recognition, demonstrate the robustness and generalization capability of the proposed framework. In Houston 2013, the method achieves 96.4% (RGB) and 97.3% (hyperspectral) overall accuracy, with cross-modality transfer performance of 87.2% (RGB → HSI) and 88.7% (HSI → RGB), further improving to 97.0% and 97.8% under joint training. On UC Merced and Indian Pines, the model attains 98.7% and 97.6% overall accuracy, respectively. These results establish semantic representation unification as a scalable and effective alternative to conventional multimodal fusion for heterogeneous remote sensing environments.
Sajid et al. (Wed,) studied this question.