Given a large un-transcribed corpus of speech utterances, we address the problem of how to select a good subset for wordlevel transcription under a given fixed transcription budget. We employ submodular active selection on a Fisher-kernel based graph over un-transcribed utterances. The selection is theoretically guaranteed to be near-optimal. Moreover, our approach is able to bootstrap without requiring any initial transcribed data, whereas traditional approaches rely heavily on the quality of an initial model trained on some labeled data. Our experiments on phone recognition show that our approach outperforms both average-case random selection and uncertainty sampling significantly.
No takes yet. Share an insight, caveat, or question.
Lin et al. (2009) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: