We present the Zero Resource Speech Challenge 2019, which proposes to build a synthesizer without any text or phonetic labels: hence, TTS without T(text-to-speech without text). We provide raw audio for a target voice in an language (the Voice dataset), but no alignment, text or labels. must discover subword units in an unsupervised way (using the Unit dataset) and align them to the voice recordings in a way that works for the purpose of synthesizing novel utterances from novel speakers, to the target speaker's voice. We describe the metrics used for, a baseline system consisting of unsupervised subword unit discovery a standard TTS system, and a topline TTS using gold phoneme. We present an overview of the 19 submitted systems from 10 and discuss the main results.
No takes yet. Share an insight, caveat, or question.
Dunbar et al. (2019) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: