Foundation models have shown remarkable generalization ability for few-shot learning (FSL). However, their potential for remote sensing scene classification has not been fully explored. Existing methods mainly adapt a single vision–language model and seldom exploit the complementary strengths of different foundation models. Moreover, the inconsistency between visual and textual representations limits the effectiveness of cross-modal learning under limited supervision. To address these issues, we propose a framework for few-shot remote sensing scene classification leveraging collaboration of foundation models. Specifically, the proposed framework integrates a large language model, a text-to-image diffusion model, and a vision–language model to enrich class semantics, synthesize category-related training samples, and learn transferable visual-textual representations. Furthermore, a Deep Cross-modal Alignment (DCA) module is developed to improve feature consistency across modalities. The DCA module incorporates multi-scale visual features, lightweight adapters, and a contrastive learning objective to obtain more discriminative task-adaptive representations. Extensive experiments on 12 remote sensing scene classification datasets under various few-shot settings demonstrate that the proposed framework achieves competitive performance compared with existing prompt learning and efficient parameter tuning methods, with consistent improvements observed on average across the evaluated datasets. Recent multimodal large language models evaluated in the zero-shot setting are additionally reported as references, while the 2-shot results of our method illustrate the performance when limited labeled data are available. This comparison provides a broader view of the trade-offs between label availability, classification accuracy, and inference cost across the two paradigms.
No takes yet. Share an insight, caveat, or question.
Ji et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: