PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 10, 2016IEEE/ACM Transactions on Audio Speech and Language Processing156 citations

A Joint Training Framework for Robust Automatic Speech Recognition

View Full Paper
ZWZhong-Qiu WangDWDeLiang Wang

Key Points

Key points are not available for this paper at this time.

Abstract

Robustness against noise and reverberation is critical for ASR systems deployed in real-world environments. In robust ASR, corrupted speech is normally enhanced using speech separation or enhancement algorithms before recognition. This paper presents a novel joint training framework for speech separation and recognition. The key idea is to concatenate a deep neural network (DNN) based speech separation frontend and a DNN-based acoustic model to build a larger neural network, and jointly adjust the weights in each module. This way, the separation frontend is able to provide enhanced speech desired by the acoustic model and the acoustic model can guide the separation frontend to produce more discriminative enhancement. In addition, we apply sequence training to the jointly trained DNN so that the linguistic information contained in the acoustic and language models can be back-propagated to influence the separation frontend at the training stage. To further improve the robustness, we add more noise- and reverberation-robust features for acoustic modeling. At the test stage, utterance-level unsupervised adaptation is performed to adapt the jointly trained network by learning a linear transformation of the input of the separation frontend. The resulting sequence-discriminative jointly-trained multistream system with run-time adaptation achieves 10.63% average word error rate (WER) on the test set of the reverberant and noisy CHiME-2 dataset (task-2), which represents the best performance on this dataset and a 22.75% error reduction over the best existing method.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2016) studied this question.

synapsesocial.com/papers/6a1cae1e3e9e446a9a85c6f5https://doi.org/10.1109/taslp.2016.2528171
Ask AI
Helpful
Bookmark
Share
View Full Paper