Speech restoration is becoming a foundational technology as it enables to use in-the-wild speech data for the generative AI training by removing artifacts. We present Sidon: a robust and high fidelity speech restoration model based on parametric speech resynthesis. Sidon is highly multilingual with 102 languages and outputs full fidelity speech in 48 kHz sampling rate. Upon training multiple artifacts are simulated such as background noise, reverberation, band limiting, codec, clipping, wind noise and packet loss. We evaluate the performance of Sidon against Google's internal Miipher and Miipher-2. The result shows competitive performance in several metrics. We also show that the Sidon is well generalized to the in-the-wild noisy speech recordings.
Nakata et al. (Wed,) studied this question.