GSViTS improves speech naturalness and fluency in parallel TTS models, suggesting advancements in efficiency and quality.
Parallel end-to-end TTS models have made notable progress in speech naturalness but still suffer from limitations in fluency and computational efficiency. This letter presents GSViTS, a parallel TTS framework that achieves robust alignment and enchanced prosody by integrating a differentiable SoftDTW-based alignment module with uncertainty-aware duration prediction. In addition, a Dual-Path Gated Linear Attention (DP-GLA) mechanism is introduced to support efficient long-sequence modeling with reduced computational overhead. Experiments on LJSpeech and VCTK demonstrate 20% MOS improvement and 7% lower computational cost over baselines.
No takes yet. Share an insight, caveat, or question.
Zhao et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: