PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026IEEE Transactions on Image Processing3 citations

Disentangle to Fuse: Towards Content Preservation and Cross-Modality Consistency for Multi-Modality Image Fusion

View Full Paper
XQXinran QinYCYuning CuiSSShangquan Sun

Key Points

  • The central aim is to enhance multi-modal image fusion by disentangling style and content representations for improved output quality.
  • Developed a framework called C<sup>2</sup>MFuse for multi-modal image fusion.
  • Implemented a content-preserving style normalization mechanism to maintain scene structure.
  • Aggregated normalized features to enhance fine-grained details in fused images.
  • Aligned fused representation with a defined source modality to increase semantic consistency.
  • Introduced an adaptive consistency loss with learnable transformation for global consistency.
  • C<sup>2</sup>MFuse achieved superior image fusion quality compared to existing methods.
  • Demonstrated effective generalization across various visual applications.
  • Performed extensive experiments on five datasets, validating high-quality performance.

Abstract

Multi-modal image fusion (MMIF) aims to integrate complementary information from heterogeneous sensor modalities. However, substantial cross-modality discrepancies hinder joint scene representation and lead to semantic degradation in the fused output. To address this limitation, we propose C2MFuse, a novel framework designed to preserve content while ensuring cross-modality consistency. To the best of our knowledge, this is the first MMIF approach to explicitly disentangle style and content representations across modalities for image fusion. C2MFuse introduces a content-preserving style normalization mechanism that suppresses modality-specific variations while maintaining the underlying scene structure. The normalized features are then progressively aggregated to enhance fine-grained details and improve content completeness. In light of the lack of ground truth and the inherent ambiguity of the fused distribution, we further align the fused representation with a well-defined source modality, thereby enhancing semantic consistency and reducing distributional uncertainty. Additionally, we introduce an adaptive consistency loss with learnable transformation, which provides dynamic, modality-aware supervision by enforcing global consistency across heterogeneous inputs. Extensive experiments on five datasets across three representative MMIF tasks demonstrate that C2MFuse achieves efficient and high-quality fusion, surpasses existing methods, and generalizes effectively to downstream visual applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qin et al. (2026) studied this question.

synapsesocial.com/papers/6980ff19c1c9540dea811cb7https://doi.org/10.1109/tip.2026.3657183
Ask AI
Helpful
Bookmark
Share
View Full Paper