Recently there has been interest in the approaches for train-ing speech recognition systems for languages with limited re-sources. Under the IARPA Babel program such resources have been provided for a range of languages to support this research area. This paper examines a particular form of approach, data augmentation, that can be applied to these situations. Data aug-mentation schemes aim to increase the quantity of data available to train the system, for example semi-supervised training, multi-lingual processing, acoustic data perturbation and speech syn-thesis. To date the majority of work has considered individual data augmentation schemes, with few consistent performance contrasts or examination of whether the schemes are comple-mentary. In this work two data augmentation schemes, semi-supervised training and vocal tract length perturbation, are ex-amined and combined on the Babel limited language pack con-figuration. Here only about 10 hours of transcribed acoustic data are available. Two languages are examined, Assamese and Zulu, which were found to be the most challenging of the Ba-bel languages released for the 2014 Evaluation. For both lan-guages consistent speech recognition performance gains can be obtained using these augmentation schemes. Furthermore the impact of these performance gains on a down-stream keyword spotting task are also described. Index Terms: data augmentation, speech recognition, babel 1.
No takes yet. Share an insight, caveat, or question.
Ragni et al. (2014) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: