Comparative evaluation demonstrates superior performance of multilingual self-supervised models across multilingual audio corpora, highlighting the advantage of cross-language pre-training.
Multilingual speech emotion recognition (SER) represents a challenging task that bridges monolingual and cross-lingual approaches, benefiting from larger datasets while leveraging in-group cultural advantages. While previous studies focused on architectural solutions, recent advances in self-supervised learning (SSL) for automatic speech recognition have produced multilingual pre-trained models with superior performance. This study systematically evaluates 13 SSL models, including five models specifically trained on multilingual data. We conducted experiments on the EmoFilm dataset containing English, Italian, and Spanish samples using both fixed split and five-fold cross-validation scenarios in speaker-independent settings. Our results demonstrate that multilingual SSL models, particularly XLS-R variants, significantly outperform previous state-of-the-art monolingual SSL models such as UniSpeech-SAT and WavLM for multilingual SER tasks. The findings suggest that multilingual pre-training provides substantial advantages for cross- and within-language emotion recognition, with XLS-R achieving the highest performance across all evaluation scenarios. This work provides the first comprehensive comparison of multilingual SSL models for SER and establishes new benchmarks for multilingual speech emotion recognition research. Additionally, we evaluated the improvements for each language as well as for each emotion category.
No takes yet. Share an insight, caveat, or question.
Atmaja et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: