Since single image from single-modality measurement cannot completely represent scene content, advanced devices based on various sensors are widely used to capture different images for multi-modality image fusion according to different mechanisms. However, traditional image fusion methods always use image-level decomposition to obtain high-frequency and low-frequency for each frequency-band combination of different modalities and multiple-band fusion, whose extractors are not learnable. On the contrary, deep-learning image fusion approaches can use learnable convolutional layers to build up an end-to-end network. Most of them simply use feature-level fusion and produce only one fusion result. To this end, we propose a diversified multi-modal recurrent-octave fusion method based on image-level and feature-level modulation for multi-modal image fusion. Specifically, our method includes Global-Coefficient Modulation (GCM) network and Recurrent-Octave Auto-encoder (ROA) network as well as HYper-Prior (HYP) network. Firstly, the input image is decomposed as the structure component and texture component by Gaussian filter operator, after which the GCM network is designed to predict weighted coefficient to combine these two components as an auxiliary modulation map for assisting ROA network. Secondly, the encoder of the ROA network is introduced to condense the input map and auxiliary modulation map as bottle-neck features, while the bottle-neck features are expanded to reconstruct the input map by its decoder. Thirdly, in the encoder, recurrent-octave convolution layer is proposed to decompose features and reassemble features, while different layers are connected in the recurrent form. Fourthly, we build an HYP network to estimate modulation coefficient, which is leveraged to update the output features from recurrent-octave convolution layer. Finally, substantial experimental results indicate that the proposed method has higher fusion performances and can produce diversified fusion maps as compared with the newest image fusion methods.
No takes yet. Share an insight, caveat, or question.
Zhang et al. (2025) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: