Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts still suffer from two types of latencies: (a) the computational latency (synthesizing time), which grows linearly with the sentence length, and (b) the input latency in scenarios where the input text is incrementally available (such as in simultaneous translation, dialog generation, and assistive technologies). To reduce these latencies, we propose a neural incremental TTS approach using the prefix-to-prefix framework from simultaneous translation. We synthesize speech in an online fashion, playing a segment of audio while generating the next, resulting in an O(1) rather than O(n) latency. Experiments on English and Chinese TTS show that our approach achieves similar speech naturalness compared to full sentence TTS, but only with a constant (1-2 words) latency. * M. M. and B. Z. contributed equally; M. M. co-directed the project (with L. H.), and was responsible for the majority of ideas and implementations; B. Z. improved the speech quality significantly and implemented the ideas on different TTS systems.
No takes yet. Share an insight, caveat, or question.
Ma et al. (2020) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: