Key points are not available for this paper at this time.
Automatic speech emotion recognition (SER) is the task to predict the emotional state from speech, and has become a vital part of human-machine interaction. Recent studies focus on learning from multiple features via feature fusion. This paper proposes a lightweight fusion model with time-frequency features for SER. First, time-frequency features are extracted from the Mel frequency cepstral coefficients (MFCC) of speech by using parallel convolutional neural networks, and enhanced by using a dimension-specific weighting method. Then, a two-stage feature fusion model is proposed to fully capture the complementary information between time and frequency features and to reduce the redundancy. It consists of an adaptive fusion module to redistribute the features for a local fusion, and a global fusion module to enhance the cross-dimension feature communication. Experimental results show that our model achieves competitive performance on the datasets of IEMOCAP and RAVDESS, with recognition accuracies of 74.62% and 86.11%, respectively. The model contains only 0.82M and 0.85M parameters for the two datasets, which is lightweight for SER tasks.
Zhang et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: