To address the semantic gap and complex feature entanglement inherent in multimodal emotion recognition, we propose the Dynamic Heterogeneous Graph Temporal Network (DHGTN), an end-to-end framework designed to model dynamic cross-modal interactions effectively. Utilizing a robust backbone of Wav2vec 2.0, VideoMAE, and BERT, we introduce a “Shared Private” subspace projection mechanism that explicitly disentangles emotion common features from modality-specific noise through contrastive learning to ensure strict semantic alignment. Furthermore, our collaborative Dynamic Heterogeneous Graph and Transformer module overcomes static fusion limitations by constructing time-varying graphs for instantaneous associations and employing global attention to capture long-range temporal dependencies. Extensive experiments on the IEMOCAP and MELD benchmarks demonstrate that DHGTN significantly outperforms state-of-the-art baselines, achieving weighted F1-scores of 73.86% and 66.87%, respectively, which confirms the method’s effectiveness and robustness.
Da et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: