PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 21, 2026Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery0 citations

Speech Emotion Recognition Using Transfer Learning: A Comparative Study

View Full Paper
YKYunus Korkmaz

Key Points

  • The aim is to evaluate transfer learning methods for enhancing speech emotion recognition accuracy and generalization.
  • Reviewed existing studies on transfer learning for speech emotion recognition.
  • Categorized techniques into feature embedding extraction and fine-tuning strategies.
  • Analyzed performance across well-known emotion datasets such as RAVDESS and EmoDB.
  • Transfer learning significantly improves recognition accuracy compared to traditional methods.
  • Fine-tuned models generally outperform fixed feature encoders.
  • Comparison reveals methodological trends and performance metrics like accuracy and F1 score.

Abstract

ABSTRACT Speech emotion recognition (SER) is a critical research area at the intersection of affective computing and audio signal processing. Traditional approaches often require extensive manual feature engineering and large labeled datasets, which can limit performance in real‐world applications. Recently, transfer learning with pretrained audio neural networks has gained significant traction for SER, leveraging knowledge from large‐scale audio corpora to improve recognition accuracy and generalization. This review presents a comprehensive overview of transfer learning‐based SER systems, with a particular focus on pretrained audio models such as YAMNet, VGGish, Wav2Vec2, and so on. Two main paradigms were examined in this study. First, feature embedding extraction, where pretrained models serve as fixed feature encoders. Second, fine‐tuning strategies, where model weights are partially or fully updated on emotion‐specific corpora. The article categorizes and compares state‐of‐the‐art studies across multiple datasets, including RAVDESS, EmoDB, IEMOCAP, CREMA‐D, and so on, while discussing commonly used evaluation metrics such as accuracy, precision, recall, F1 score, and AUC. Comparative tables are presented to highlight methodological trends and performance differences between feature‐based and fine‐tuned SER systems. This overview aims to guide researchers in selecting appropriate models and strategies for future SER tasks and to identify open challenges and future research directions in transfer learning for emotion recognition from speech signals. This article is categorized under: Technologies > Machine Learning Technologies > Artificial Intelligence Algorithmic Development > Multimedia

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yunus Korkmaz (2026) studied this question.

synapsesocial.com/papers/69994bdd873532290d01ff41https://doi.org/10.1002/widm.70073
Ask AI
Helpful
Bookmark
Share
View Full Paper