PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 20156,096 citations

Librispeech: An ASR corpus based on public domain audio books

View Full Paper
VPVassil PanayotovGCGuoguo ChenDPDaniel Povey

Key Points

  • To introduce and evaluate LibriSpeech, a large-scale public domain read English speech corpus designed for training and benchmarking automatic speech recognition systems.
  • Extracted and prepared 1,000 hours of read English speech sampled at 16 kHz derived from LibriVox public domain audiobooks.
  • Constructed complementary language-model training data, pre-built language models, and Kaldi training scripts for open distribution.
  • Trained acoustic models on LibriSpeech and evaluated their performance against standard Wall Street Journal (WSJ) test sets.
  • Acoustic models trained on LibriSpeech achieved lower word error rates on WSJ test sets than models trained directly on the WSJ corpus.
  • Released a publicly accessible 1,000-hour speech recognition benchmark with accompanying language models and baseline recipes.

Abstract

This paper introduces a new corpus of read English speech, suitable for training and evaluating speech recognition systems. The LibriSpeech corpus is derived from audiobooks that are part of the LibriVox project, and contains 1000 hours of speech sampled at 16 kHz. We have made the corpus freely available for download, along with separately prepared language-model training data and pre-built language models. We show that acoustic models trained on LibriSpeech give lower error rate on the Wall Street Journal (WSJ) test sets than models trained on WSJ itself. We are also releasing Kaldi scripts that make it easy to build these systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Panayotov et al. (2015) studied this question.

synapsesocial.com/papers/69d72b1166e6af6209f507f4https://doi.org/10.1109/icassp.2015.7178964
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1An Empirical Study of Smoothing Techniques for Language Modeling1996 · 632 citations
  2. 2Boosted MMI for model and feature-space discriminative training2008 · 348 citations
  3. 3Lightly supervised recognition for automatic alignment of large coherent speech recordings2010 · 101 citations
  4. 4The design for the wall street journal-based CSR corpus1992 · 1,140 citations
  5. 5Identification of common molecular subsequences1981 · 10,145 citations