PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 20250 citationsOpen Access

DCA-CL: Enhancing Multimodal Emotion Recognition via Dual Cross Attention and Contrastive Learning

View Full Paper
XWXin WangSLShubo LiuHDHongshe Dang

Key Points

  • The DCA-CL framework significantly enhances emotion recognition accuracy, especially in few-shot settings.
  • Experiments on the IEMOCAP and MELD datasets validate that the model effectively integrates cross-modal information.
  • Incorporation of dynamic focal contrastive loss boosts the model's ability to learn discriminative representations.
  • The framework utilizes bidirectional cross-modal attention and temporal gating to filter salient features.

Abstract

Abstract Emotion is a subjective human response to external events or stimuli and plays a crucial role across various application domains. Consequently, emotion recognition has become a central focus of research. However, existing mainstream approaches still face several challenges, such as limited interaction across different modalities and low recognition accuracy when dealing with limited samples involving semantically similar but categorically distinct emotions. To tackle these challenges, we introduce a new multimodal emotion recognition framework, named DCA-CL (Dual Cross Attention with Contrastive Learning), which aims to improve the integration and effectiveness of cross-modal information. The proposed model incorporates a feature fusion network that combines bidirectional cross-modal attention with self-attention mechanisms, enabling effective modeling of both intra-modal and cross-modal interactions. Furthermore, a temporal gating mechanism is adopted to filter salient features and suppress redundant information, while dynamic weight allocation facilitates efficient fusion of modality-specific features. During the training phase, a dynamic modal distillation mechanism is introduced to dynamically select the optimal teacher mode based on modal quality, guiding weak modes to learn high-quality semantic features and enhance their ability to represent weak modes;To enhance recognition accuracy in few-shot settings and among semantically close emotion categories, we incorporate a dynamic focal contrastive loss, which boosts the model’s ability to learn discriminative representations. Experiments conducted on the IEMOCAP and MELD datasets confirm that the proposed DCA-CL framework delivers outstanding overall performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68bb3a492b87ece8dc9557a3https://doi.org/10.21203/rs.3.rs-7251362/v1
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CLDAE: A Two Stage EEG-Based Emotion Recognition Framework Combining Contrastive Learning and Dual-Attention Encoder2026
  2. 2Multimodal emotion recognition based on multi-head cross-attention mechanism2025
  3. 3A modality-aware contrastive learning framework for multimodal sentiment analysis2026
  4. 4Hypercomplex Neural Network and Cross-Modal Attention for Multi-Modal Emotion Recognition Using Physiological Signals2025 · 10 citations
  5. 5Self-Learning Multimodal Emotion Recognition Based on Multi-Scale Dilated Attention2026