Analysis demonstrates differences in vocal emotions between languages, indicating the role of emotional prosody.
This study examined the acoustic profiles of five basic emotions in American English and Mandarin Chinese using a big data approach. A total of 6,373 features were extracted using the openSMILE toolkit, and key discriminative features were identified through random forest classification. In American English, vocal emotions were primarily conveyed through pitch-related features, while Mandarin Chinese, shaped by its tonal constraints, relied more on spectral and voice quality cues, including MFCCs, HNR, and shimmer. Linear mixed-effects models confirmed significant effects of emotion on the top-ranked features, and Cohen’s d further supported distinct acoustic profiles for each emotion. K-means clustering revealed both categorical groupings and dimensional overlaps, such as the clustering of high-arousal emotions like happy and surprised, and low-arousal emotions like sad and neutral. These results suggest that vocal emotion expression is shaped by language-specific prosodic systems, as well as by both discrete emotion categories and continuous affective dimensions, supporting an integrated model of emotional prosody.
No takes yet. Share an insight, caveat, or question.
Fenqi Wang (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: