PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 1, 2013555 citations

Ideal ratio mask estimation using deep neural networks for robust speech recognition

View Full Paper
ANArun NarayananDWDeLiang Wang

Key Points

Key points are not available for this paper at this time.

Abstract

We propose a feature enhancement algorithm to improve robust automatic speech recognition (ASR). The algorithm estimates a smoothed ideal ratio mask (IRM) in the Mel frequency domain using deep neural networks and a set of time-frequency unit level features that has previously been used to estimate the ideal binary mask. The estimated IRM is used to filter out noise from a noisy Mel spectrogram before performing cepstral feature extraction for ASR. On the noisy subset of the Aurora-4 robust ASR corpus, the proposed enhancement obtains a relative improvement of over 38% in terms of word error rates using ASR models trained in clean conditions, and an improvement of over 14% when the models are trained using the multi-condition training data. In terms of instantaneous SNR estimation performance, the proposed system obtains a mean absolute error of less than 4 dB in most frequency channels.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Narayanan et al. (2013) studied this question.

synapsesocial.com/papers/6a16bc0fb082e78ad77b899ahttps://doi.org/10.1109/icassp.2013.6639038
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Computational Auditory Scene Analysis2006 · 517 citations
  2. 2Techniques for Noise Robustness in Automatic Speech Recognition2012 · 119 citations
  3. 3Speech segregation based on sound localization2003 · 399 citations
  4. 4A Fast Learning Algorithm for Deep Belief Nets2006 · 16,549 citations
  5. 5DARPA TIMIT:1993 · 1,292 citations