Review demonstrates transition from rule-based pipelines to deep neural speech synthesis, highlighting emerging paradigms in diffusion models and voice cloning.
Text-to-Speech synthesis has evolved from rule-based and concatenative systems to highly natural neural architectures. This review surveys the historical development of TTS, core system components, representative model families, benchmark datasets, and evaluation protocols. We summarize major advances in end-to-end neural TTS, non-autoregressive generation, neural vocoders, multilingual and expressive synthesis, and zero-shot voice cloning. We additionally discuss persistent challenges related to prosody control, low-resource languages, real-time inference, and ethical risks. Finally, we outline promising research directions in diffusion-based generation and speech-language foundation models for controllable and robust TTS.
No takes yet. Share an insight, caveat, or question.
Do et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: