This work proposes the usage of multimodal large language models for speech emotion recognition (SER) on low-resource settings for Iberian languages. Given the existing low amount of annotated SER data for other languages, we also propose a pipeline for generating high-quality synthetic data using existing emotional text-to-speech (TTS) and their cloning capabilities. Specifically, we design a selective, quality-controlled TTS pipeline combining LLM-ensemble translation with self-verification and expressive voice cloning, followed by automatic ASR-WER, speaker-similarity, and emotion filters. This approach introduces a novel filtering strategy that ensures synthetic data reliability. The resulting data include MSP-MEA, 1 1 https://huggingface.co/datasets/jaimebellver/SER-MSPMEA-Spanish . a synthetic Spanish extension of MSP-Podcast. Building on our previous multimodal SER framework, we compare the usage of frozen LLM as classifiers with an MLP baseline and evaluate classical versus TTS-based augmentation across five corpora (IEMOCAP, MEACorpus, EMS, VERBO, AhoEmo3). The best configuration (W2v-BERT-2 → attentive pooling → frozen Bloomz-7b1) improves mean F1 by +4.9 point increase over a MLP head baseline. Among augmentation techniques, Mix-up remains the most robust overall, while TTS achieves competitive performance, surpassing traditional data augmentation techniques on EMS and VERBO. These results indicate that carefully filtered TTS data can complement classical perturbations, providing a viable, dataset-dependent strategy for multilingual SER. Code, models, and datasets are publicly released. 2 2 https://github.com/jaimebs2/SpeechFactory . • Propose SER with frozen LLMs over acoustic embeddings on 5 multilingual corpora. • Release a quality-controlled TTS data augmentation pipeline for emotional speech. • Show that TTS-based augmentation can match or surpass classical methods. • Release trained models and MSP-MEA, the first synthetic Spanish extension of MSP-Podcast.
No takes yet. Share an insight, caveat, or question.
Bellver-Soler et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: