PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20246 citationsOpen Access

Inter-Modality and Intra-Sample Alignment for Multi-Modal Emotion Recognition

View Full Paper
YWYusong WangDLDongyuan LiJSJialun Shen

Key Points

Key points are not available for this paper at this time.

Abstract

Multi-modal emotion recognition aims to recognize different emotions precisely by utilizing different modality information. Current multi-modal methods always suffer the pain of the modality gap. Besides, many works use attention mechanisms to fuse different modalities features, which reduces the diversity of the modalities. To deal with these issues, we propose an inter-modality and intra-sample alignment framework for multi-modal emotion recognition. Specifically, we first map text features and audio features to a shared subspace and align them using a cross-modal self-attention mechanism and Maximum Mean Discrepancy matrix. To preserve the diversity of each modality, we extract modality-specific features using different encoders. Finally, we employ supervised contrastive learning techniques to align the features at a sample-level. Extensive experiments indicate that our method achieves state-of-the-art performance and effectiveness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2024) studied this question.

synapsesocial.com/papers/68e7376bb6db6435876b0fe9https://doi.org/10.1109/icassp48485.2024.10446571
Ask AI
Helpful
Bookmark
Share
View Full Paper