PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 14, 2017IEEE Journal of Selected Topics in Signal Processing10 citationsOpen Access

Inverted Alignments for End-to-End Automatic Speech Recognition

View Full Paper
PDPatrick DoetschMHMirko HannemannRSRalf Schlüter

Key Points

Key points are not available for this paper at this time.

Abstract

In this paper, we propose an inverted alignment approach for sequence classification systems like automatic speech recognition (ASR) that naturally incorporates discriminative, artificial-neural-network-based label distributions. Instead of aligning each input frame to a state label as in the standard hidden Markov model (HMM) derivation, we propose to inversely align each element of an HMM state label sequence to a segment-wise encoding of several consecutive input frames. This enables an integrated discriminative model that can be trained end-to-end from scratch or starting from an existing alignment path. The approach does not assume the usual decomposition into a separate (generative) acoustic model and a language model, and allows for a variety of model assumptions, including statistical variants of attention. Following our initial paper with proof-of-concept experiments on handwriting recognition, the focus of this paper was the investigation of integrated training and an inverted decoding approach, whereas the acoustic modeling still remains largely similar to standard hybrid modeling. We provide experiments on the CHiME-4 noisy ASR task. Our results show that we can reach competitive results with inverted alignment and decoding strategies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Doetsch et al. (2017) studied this question.

synapsesocial.com/papers/6a16eed57cba52b0f77bbccchttps://doi.org/10.1109/jstsp.2017.2752691
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Linear discriminant analysis for improved large vocabulary continuous speech recognition1992 · 321 citations
  2. 2Classification and recognition with direct segment models2012 · 19 citations
  3. 3The blame game in meeting room ASR: An analysis of feature versus model errors in noisy and mismatched conditions2013 · 15 citations