Real-world speech restoration must handle coupled distortions, including acoustic noise and reverberation, codec artifacts, clipping, and artifacts left by upstream enhancement systems. Token-based generative systems offer a flexible route for such universal restoration, but discrete audio tokens can discard fine acoustic detail, and aggressive generative decoding may over-process inputs that are already close to clean speech. We propose AudioVAE-MASR, a continuous-latent masked autoregressive framework for multi-distortion speech restoration. A frozen AudioVAE maps clean and degraded speech into paired continuous latent sequences; a Conformer-based branch extracts the degraded-condition sequence Cy from degraded latents; a two-stream masked autoregressive encoder-decoder conditions masked clean-latent recovery on both degraded context and visible clean tokens; and a lightweight diffusion head models the masked clean tokens in the continuous latent space. On the released CCF AATC 2025 blind test set, the main inference setting (K=16, temperature 0.5) achieved WAcc 0.793, SIG 3.401, BAK 3.987, OVRL 3.111, PESQ 1.780, and ESTOI 0.798. Relative to the degraded input, these results improved WAcc and DNSMOS but did not improve PESQ; relative to the organizer baseline, they improved WAcc, SIG, OVRL, and PESQ but remained lower in BAK. A local subjective MOS evaluation with five listeners gave an overall mean score of 4.08 for AudioVAE-MASR, compared with 3.70 for the degraded input and 4.59 for the clean reference. Distortion-type, ablation, and parameter-sensitivity analyses further show that codec inputs remain vulnerable to over-restoration and that longer iterative decoding does not provide a consistent gain. The study therefore presents AudioVAE-MASR as a transparent continuous-latent restoration framework and identifies the fidelity-control problems that must be solved before such generative restoration can match the strongest lightweight discriminative systems.
Hu et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: