This study proposes an effective angry speech detection approach by leveraging content structure within the input speech. A classifier based on an “emotional” language model score is formulated and combined with acoustic feature based classifiers including TEO-based feature and conventional Mel frequency cepstral coefficients (MFCC). The proposed detection algorithm is evaluated on real-life conversational speech which was recorded between customers and call center operators over a telephone network. Analysis on the conversational speech corpus presents a distinctive property between neutral and angry speech in word distribution and frequently occurring words. An improvement of up to 6.23% in Equal Error Rate (EER) is obtained by combining the TEO-based and MFCC features, and emotional language model score based classifiers.
No takes yet. Share an insight, caveat, or question.
Kim et al. (2010) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: