Training text-to-speech (TTS) systems on project-collected real-world recordings collected under multiple recording conditions that exhibit both environmental noise and atypical prosody is challenging, especially for tonal languages where prosodic structure directly affects intelligibility. This study adopts a multi-objective perspective on data curation that balances spectral fidelity and prosodic integrity and instantiates this perspective on a 24 hour screened subset drawn from a real-world Zhuang corpus of approximately 100 hours. Seven preprocessing strategies were compared under a controlled setting in which the training roster and model architecture were held fixed. The main empirical finding is that disabling internal silence compression while retaining trimming, filtering, and normalization (noᵢntₛil) yielded a substantial perceptual gain: Mean Opinion Scores (MOS) for overall naturalness improved from 3. 17 for a minimally processed baseline to 3. 72 (p < 0. 05), even though a heavily cleaned configuration (fullₘethod) attained the lowest Fréchet Audio Distance (FAD) among all systems. This pattern indicates that, in this project-collected Wuming Zhuang setting and under a VITS backbone, aggressive internal silence compression can act as a prosody-distorting operation: it improves apparent spectral cleanliness but perturbs boundary-related timing and pitch behavior that listeners rely on for naturalness.
Bai et al. (Fri,) studied this question.