Deep learning evaluation demonstrates superior resolution and spectral fidelity in remote sensing image fusion, indicating better feature optimization via hybrid Swin Transformer architectures.
Multispectral and panchromatic image fusion technology (also known as pan-sharpening) is essential for generating high-resolution multispectral remote sensing imagery, yet existing deep learning-based methods often struggle to effectively balance spatial detail enhancement with spectral fidelity and lack sufficient feature interaction. Therefore, to address these issues, a lightweight tri-branch convolutional Swin Transformer network with hybrid feature injection for pan-sharpening, i.e. TCSwinPNet, is proposed in this paper. In this novel architecture, the hybrid branch employs Swin Transformer as the backbone to synchronously inject hybrid features into the panchromatic and multispectral branches, thereby achieving cross-modal interaction. To address the two critical requirements of spatial detail enhancement and spectral fidelity, the lightweight spatial detail perception enhancement module and the spectral band collaborative preservation module are designed. Furthermore, attention mechanisms based on the Kolmogorov-Arnold Network (KAN) are introduced to utilize the composition of nested nonlinear functions to adaptively model the complex spatial–spectral coupling relationships in remote sensing imagery, enabling more refined feature optimization. Comprehensive experimental results demonstrate that TCSwinPNet consistently outperforms state-of-the-art methods in qualitative results, quantitative metrics, and land-use classification tasks. https://github.com/RSIDEA-ECUT/TCSwinPNet.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: