Multimodal cross-domain few-shot learning has advanced hyperspectral image (HSI) classification by introducing textual semantics as auxiliary supervision to guide visual representation learning and improve the discriminability of class prototypes under limited labeled samples. However, existing methods often rely on coarse or static prompts, which makes it difficult to capture complex spatial structures, subtle spectral responses, and domain-varying spatial-spectral patterns. To address this issue, we propose a visually guided dual-level semantic alignment framework (VGDLA) for cross-domain few-shot HSI classification. Specifically, we design a visually guided fine-grained prompt generator (VFPG) to map spatial and spectral representations from the visual encoder into prompt tokens, enabling the text encoder to generate task-adaptive visually conditioned semantics. A dual-scale semantic fusion module (DS-SFM) then integrates these adaptive semantics with stable template-based class priors to construct reliable semantic representations. Furthermore, we develop a dual-level semantic contrastive learning strategy to overcome the limitations of single semantic supervision, which cannot simultaneously ensure fine-grained semantic discrimination and stable category-level consistency. The proposed strategy employs fused visual-conditioned semantics for fine-grained prototype alignment and template-based semantics for class-level regularization. Experiments on four benchmark HSI datasets demonstrate that VGDLA consistently outperforms state-of-the-art cross-domain few-shot HSI classification methods.
No takes yet. Share an insight, caveat, or question.
Lin et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: