Hyperspectral–multispectral (HSI–MSI) image fusion aims to reconstruct high-spatial-resolution hyperspectral images (HR-HSIs) by combining the spectral fidelity of low-resolution HSIs (LR-HSIs) with the spatial details of high-resolution MSIs (HR-MSIs). A key challenge is preserving spectral–spatial consistency under cross-modal resolution mismatch, where inadequate long-range dependency modeling and unstable inter-modality interaction may induce spectral distortion and structural discontinuities. This paper proposes DSIR-Net (DSIR), a dual-stream state-space fusion architecture equipped with an implicit neural representation (INR) module. DSIR decouples spectral and spatial representation learning into two coordinated streams and leverages state-space modeling to aggregate global context efficiently during progressive fusion. Moreover, INR-based coordinate-conditioned refinement provides continuous sub-pixel compensation, enhancing high-frequency detail recovery while suppressing fusion-induced artifacts. Across four commonly used benchmark datasets, DSIR shows consistent advantages over the competing methods in both numerical metrics and visual reconstruction quality. In addition to sharper structural details, DSIR preserves spectral information more faithfully. Using the best result among the baselines on each dataset as reference, the PSNR improvements are 0.040 dB (Houston), 0.204 dB (PaviaU), 0.093 dB (Botswana), and 0.163 dB (Chikusei).
Liu et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: