Deep learning evaluation demonstrates enhanced emotion classification accuracy across multimodal data, suggesting superior affective computing performance compared to unimodal models.
Developing a Human-Computer integration system that is responsive and accurate is a challenge, particularly in multimodal emotion detection. Therefore, this research presents a dynamic multimodal deep learning framework to classify five emotion categories (angry, excited, frustrated, neutral, and sad) using the IEMOCAP dataset. This study leverages three data sources—audio, text, and images, integrates the audio modality via linguistic transcription to optimize fusion. A hybrid enrichment scheme is proposed, that combines RoBERTa, FastText, and NRC Emotion Lexicon. ResNet-18 for image feature extraction. All multimodal features are integrated using feature fusion techniques and classified with MLP, showing classification accuracy results of 86.48%, an increase of approximately 7.48% compared to the best unimodal model. This configuration achieved a Macro F1-Score of 0.8690, ensuring the system’s effectiveness across imbalanced emotional categories. To ensure objectivity in evaluating imbalanced categories, this study used macro-averaging to assess model performance. Furthermore, the superiority of the multimodal model architecture over the baseline model has been confirmed using a Z-value of 16.94 ( \(p < 0.05\) ). These results indicate that the integration of these modalities provides a real and statistically significant contribution. This confirms that the integration of three modalities can recognize human emotions accurately and effectively in mitigating unimodal limitations. This study proposes an efficient solution for the future development of multimodal affective computing systems.
No takes yet. Share an insight, caveat, or question.
Setiawan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: