A high-resolution (HR) hyperspectral imaging at video rates can be achieved by fusing a multi-band low-resolution (LR) mosaiced image and a single-band HR panchromatic (PAN) image in a single shot within a short time. However, the current fusion methods always suffer from a spatial or spectral distortion. To alleviate this, we propose an equivariant Bayesian variational inference framework. Specifically, we decompose an HR hyperspectral image (HSI) into the principal component and sparsity residual, which are modeled as latent variables with Gaussian priors. Each component is estimated via a shared deep neural network (DNN) under a variational inference framework, leveraging the shared spatial structures to enhance parameter efficiency and reconstruction accuracy. Additionally, to tackle the challenge of unavailable ground truth in real-world scenarios, we integrate the equivariant imaging (EI) prior with the Bayesian framework. By enforcing the consistency between the transformed fusion result and the re-inference output, this strategy enables the network to learn beyond the range space. Furthermore, we propose to utilize the learnable degradation functions derived from the physical imaging model to enable the proposed framework, which ensures an enhanced performance by posing plausible constraints on parameters of the degradation functions. Specifically, we explicitly model the point spread function (PSF) and spectral response function (SRF) with learnable parameters and impose non-negativity and sum-to-one constraints. Extensive experiments conducted on both simulated and real-world datasets demonstrate the effectiveness of the proposed framework, paving the way for HSI computational imaging.
Dian et al. (Thu,) studied this question.