Zero-shot speech enhancement (SE) aims to improve speech quality in unseen acoustic conditions without requiring task-specific fine-tuning. This work proposes expanded noise modeling for scalable and adaptive zero-shot speech enhancement (EN-AZS), a zero-shot SE framework with expanded noise modeling, built upon an optimized Undiff-based architecture. By extending the noise model to cover a wider range of acoustic variability and incorporating mechanisms such as speech quality scoring and coefficient calculation with controlled update strategies, EN-AZS effectively enhances speech clarity while avoiding overfitting to observed mixtures. Extensive experiments on TIMIT-N6/N9/N15, VCTK-DM and MUSAN data sets demonstrate that EN-AZS consistently outperforms both supervised and unsupervised baseline methods, particularly under mismatched noise types, speech characteristics and SNR conditions. Ablation studies further validate the importance of the speech quality guidance and coefficient calculation mechanisms and update strategies. EN-AZS provides a scalable, plug-and-play solution for robust zero-shot SE, offering a promising approach for real-world applications with diverse and unpredictable noise conditions.
Chu et al. (Thu,) studied this question.