Computational benchmark demonstrates high classification accuracy in few-shot hyperspectral datasets, indicating the utility of hybrid convolution-Mamba diffusion representations.
Key Points
To develop a linear-time generative framework that captures local spectral–spatial structure and long-range dependencies for few-shot hyperspectral image classification.
Constructed FS-MambaDiff, a two-stage hybrid framework combining local 3D convolutions with selective state-space blocks (Mamba) in a diffusion denoiser backbone.
Pretrained the diffusion model without labels, captured seven encoder–middle–decoder feature taps, and froze the denoiser to train a Mamba classifier using 25 or 30 labeled pixels per class.
Evaluated performance across 10 seeded few-shot splits on the Houston 2013, Pavia University, and WHU-Hi-HongHu benchmark datasets.
FS-MambaDiff achieved an overall accuracy of 96.02% ± 0.71% on Houston 2013, 98.53% ± 1.06% on Pavia University, and 95.11% ± 0.63% on WHU-Hi-HongHu.
Optimal diffusion feature timesteps were dataset-dependent within a low-noise regime (Houston t = 0, Pavia University t = 3, and WHU-Hi-HongHu t = 25) rather than strictly t = 0.