PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 2026IEEE Transactions on Pattern Analysis and Machine Intelligence2 citations

Decoupled Hierarchical Distillation for Multimodal Emotion Recognition

YLYong LiYWYuanzhi WangYDYi Ding

Key Points

  • The aim is to improve multimodal emotion recognition by addressing challenges from different data sources.
  • Proposed Decoupled Hierarchical Multimodal Distillation (DHMD) framework.
  • Implemented a two-stage knowledge distillation approach.
  • Used Graph Distillation Unit (GD-Unit) for coarse-grained distillation.
  • Applied a cross-modal dictionary matching for fine-grained detail alignment.
  • DHMD outperforms state-of-the-art methods in emotion recognition.
  • Achieved relative improvements of 1.3%/2.4% in accuracy on CMU-MOSI.
  • Achieved 1.3%/1.9% improvements in accuracy on CMU-MOSEI.
  • Registered 1.9%/1.8% enhancements in F1 score across datasets.

Abstract

Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1. 3%/2. 4% (ACC₇), 1. 3%/1. 9% (ACC₂) and 1. 9%/1. 8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/698584f98f7c464f2300841fhttps://doi.org/10.1109/tpami.2026.3660754
Ask AI
Helpful
Bookmark
Share
View Full Paper