PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 19, 2009IEEE Transactions on Audio Speech and Language Processing280 citations

Analysis of Emotionally Salient Aspects of Fundamental Frequency for Emotion Detection

View Full Paper
CBCarlos BussoSLSungbok LeeSNShrikanth Narayanan

Key Points

Key points are not available for this paper at this time.

Abstract

During expressive speech, the voice is enriched to convey not only the intended semantic message but also the emotional state of the speaker. The pitch contour is one of the important properties of speech that is affected by this emotional modulation. Although pitch features have been commonly used to recognize emotions, it is not clear what aspects of the pitch contour are the most emotionally salient. This paper presents an analysis of the statistics derived from the pitch contour. First, pitch features derived from emotional speech samples are compared with the ones derived from neutral speech, by using symmetric Kullback-Leibler distance. Then, the emotionally discriminative power of the pitch features is quantified by comparing nested logistic regression models. The results indicate that gross pitch contour statistics such as mean, maximum, minimum, and range are more emotionally prominent than features describing the pitch shape. Also, analyzing the pitch statistics at the utterance level is found to be more accurate and robust than analyzing the pitch statistics for shorter speech regions (e.g., voiced segments). Finally, the best features are selected to build a binary emotion detection system for distinguishing between emotional versus neutral speech. A new two-step approach is proposed. In the first step, reference models for the pitch features are trained with neutral speech, and the input features are contrasted with the neutral model. In the second step, a fitness measure is used to assess whether the input speech is similar to, in the case of neutral speech, or different from, in the case of emotional speech, the reference models. The proposed approach is tested with four acted emotional databases spanning different emotional categories, recording settings, speakers and languages. The results show that the recognition accuracy of the system is over 77% just with the pitch features (baseline 50%). When compared to conventional classification schemes, the proposed approach performs better in terms of both accuracy and robustness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Busso et al. (2009) studied this question.

synapsesocial.com/papers/6a11ea55997792fb8c8e1b5ahttps://doi.org/10.1109/tasl.2008.2009578
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Methods in Empirical Prosody Research2006 · 162 citations
  2. 2Affective Computing1997 · 5,063 citations
  3. 3Prosody beyond Fundamental Frequency2012 · 13 citations
  4. 4F0-CONTOURS IN EMOTIONAL SPEECH1999 · 50 citations
  5. 5Elements of Information Theory2001 · 38,036 citations