The proposed CH-Net framework improves multi-modal emotion recognition using physiological signals by addressing cross-modal information sharing and fusion feature representations.
Multi-modal emotion recognition plays a crucial role in human-computer interaction. Nowadays, many studies have developed fusion algorithms for this purpose. However, two challenges are still present, i.e., insufficient cross-modal information sharing and weak fusion feature representations. To this end, we develop a novel framework, namely CH-Net, for multi-modal emotion recognition with physiological signals. It is based on cross-modal attention and hypercomplex domain fusion. First, our learnable cross-modal attention mechanism adaptively aligns features across modalities, enhancing both complementarity and modality-specific discrepancies. Second, a hypercomplex fusion module encodes these features, yielding more robust representations while reducing parameter overhead. Two benchmark datasets, i.e., MAHNOB-HCI and DEAP, are utilized to train and test our model. Extensive experiments demonstrate that CH-Net is effective and outperforms state-of-the-art (SOTA) methods. Our code will be available athttps://github.com/xuxusky/CH-Net.
Xu et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: