PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 20242 citations

Impact of Heterogeneous Spectral Features for enhanced low-resource Speech Recognition System under mismatched conditions

View Full Paper
PBPuneet BawaVKVirender KadyanAMArchana Mantri

Key Points

Key points are not available for this paper at this time.

Abstract

The development of an Automatic Speech Recognition (ASR) system for children has been a significant difficulty because of the substantial inherent heterogeneity in the physical traits, articulation patterns, and mannerisms shown by each individual child. Moreover, the limited availability of substantial quantities of children's speech data may be linked to variances in vocal-tract geometries resulting from anatomical and physiological factors. The present study aims to address the aforementioned issues by conducting a study into the advancement of a voice recognition system specifically designed for children with limited resources. This study utilizes novel methods for extracting heterogeneous features from an input audio signal, which are based on raw as well as central moments. In order to mitigate the problem of limited data availability, this study utilizes different training systems that are developed using perturbation methods. Additionally, the optimization of modeling parameters is done in order to enhance the effectiveness of these models. The findings of these efforts demonstrate a significant improvement in the performance of the system. The use of a hybrid system based on a Deep Neural Network-Hidden Markov Model (DNN-HMM) on fused front end features results in a Relative Improvement of 21.36% compared to other baseline systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bawa et al. (2024) studied this question.

synapsesocial.com/papers/68e73091b6db6435876a9ef3https://doi.org/10.1109/spin60856.2024.10512314
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Cross-Accent Intelligibility of Speech in Noise: Long-Term Familiarity and Short-Term Familiarization2013 · 28 citations
  2. 2Filterbank Analysis of MFCC Feature Extraction in Robust Children Speech Recognition2019 · 9 citations
  3. 3At Home with Alexa: A Tale of Two Conversational Agents2020 · 10 citations
  4. 4Noise Shield for Microphones Used in Noisy Locations1958 · 2 citations
  5. 5Suprasegmental Features Are Not Acquired Early: Perception and Production of Monosyllabic Cantonese Lexical Tones in 4- to 6-Year-Old Preschool Children2018 · 24 citations