We investigate whether statistical regularities in discrete speech token sequences emerge before perceptual quality is achieved in speech synthesis. Using a Neural Audio Codec (NAC), we extract discrete token sequences from both natural human speech and synthetic speech generated at various training checkpoints of text-to-speech (TTS) models. Interestingly, NAC token distributions in synthetic speech resemble those of natural speech, even at early stages when outputs are unintelligible or babble-like. Across all training checkpoints, the tokens follow Zipf’s law consistently, while Heaps’ law shows milder but stable trends. These findings suggest that latent linguistic structures may begin to emerge in symbolic representations of synthetic speech, even before the models produce perceptually natural or intelligible outputs. This token-level stability contrasts with significant perceptual improvements observed across training stages, indicating that NAC tokens can reflect early structural convergence. We propose that such statistical regularities in NAC tokens offer a novel signal for early-stage model analysis or speech authenticity inspection, prior to conventional semantic or acoustic evaluations. Our work highlights the value of symbolic token analysis in understanding the internal dynamics of speech synthesis models and points toward lightweight, interpretable evaluation tools grounded in linguistic statistics.
Park et al. (Wed,) studied this question.