PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 3, 2024Automatika1 citationsOpen Access

Data augmentation using a 1D-CNN model with MFCC/MFMC features for speech emotion recognition

View Full Paper
TFT. Mary Little FlowerTJThirasama JayaSSS. Christopher Ezhil Singh

Key Points

Key points are not available for this paper at this time.

Abstract

Speech emotion recognition (SER) is attractive in several domains, such as automated translation, call centres, intelligent healthcare, and human–computer interaction. Deep learning models for emotion identification need considerable labelled data, which is only sometimes available in the SER industry. A database needs enough speech samples, good features, and a better classifier to identify emotions efficiently. This study uses data augmentation to enhance the amount of input voice samples and address the data shortage issue. The database capacity increases by adding white noise to the speech signals by data augmentation. In this work, the Mel-frequency Cepstral Coefficient (MFCC) and Mel-frequency Magnitude Coefficient (MFMC) features, along with a one-dimensional convolutional neural network (1D-CNN), are used to classify speech emotions. The datasets utilized to estimate the model's enactment were AESDD, CAFE, EmoDB, IEMOCAP, and MESD. The data augmentation with the 1D-CNN (MFMC) model performed best, with an average accuracy of 99.2% for AESDD, 99.5% for CAFE, 97.5% for EmoDB, 92.4% for IEMOCAP and 96.9% for the MESD database. The proposed 1D-CNN (MFMC) with data augmentation outperforms the 1D-CNN (MFCC) without data augmentation in emotion recognition.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Flower et al. (2024) studied this question.

synapsesocial.com/papers/68e6180bb6db6435875aad51https://doi.org/10.1080/00051144.2024.2371249
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Speech emotion recognition via graph-based representations2024 · 37 citations
  2. 2Mexican Emotional Speech Database Based on Semantic, Frequency, Familiarity, Concreteness, and Cultural Shaping of Affective Prosody2021 · 17 citations
  3. 3A database of German emotional speech2005 · 2,281 citations
  4. 4The Mexican Emotional Speech Database (MESD): elaboration and assessment based on machine learning2021 · 15 citations
  5. 5Improving Speech Emotion Recognition Using Graph Attentive Bi-Directional Gated Recurrent Unit Network2020 · 28 citations