Key points are not available for this paper at this time.
Abstract This study examined the acoustic profiles of five basic emotions in American English and Mandarin Chinese using a big data approach. A total of 6,373 features were extracted using the openSMILE toolkit, and key discriminative features were identified through random forest classification. In American English, vocal emotions were primarily conveyed through pitch-related features, while Mandarin Chinese, shaped by its tonal constraints, relied more on spectral and voice quality cues, including MFCCs, HNR, and shimmer. Linear mixed-effects models confirmed significant effects of emotion on the top-ranked features, and Cohen’s d further supported distinct acoustic profiles for each emotion. K-means clustering revealed both categorical groupings and dimensional overlaps, such as the clustering of high-arousal emotions like happy and surprised, and low-arousal emotions like sad and neutral. These results suggest that vocal emotion expression is shaped by language-specific prosodic systems, as well as by both discrete emotion categories and continuous affective dimensions, supporting an integrated model of emotional prosody.
Fenqi Wang (Fri,) studied this question.