Experimental evaluation demonstrates improved speech quality and intelligibility across noisy audio datasets, indicating effective denoising with high computational efficiency.
Speech enhancement plays a crucial role in enhancing the perceptual quality and intelligibility of speech signals that are degraded by noise. Conventional U‐Net‐based architectures effectively capture local spectral patterns but exhibit limited long‐range dependency modeling and may propagate residual noise through skip connections. Transformer‐based approaches enhance global context modeling but often incur high computational cost and insufficient preservation of fine‐grained spectral cues, limiting real‐time applicability. To address these limitations, this paper proposes a novel encoder–decoder speech enhancement framework that integrates Multi‐Scale Feature Extraction (MSFE), Dual‐Path Higher‐Order Information Interaction with Time‐Frequency Attention Module (DPH‐TFA), and Bottleneck‐Guided Feature Calibration (FC) strategy, with its hierarchical extension, Hybrid Cross‐Scale Feature Calibration (H‐CS‐FC). The MSFE blocks extract rich local patterns across multiple receptive fields, capturing both fine‐grained and global time‐frequency cues. While stacked DPH‐TFA blocks at the bottleneck model structured long‐range dependencies along time and frequency axes. The FC and H‐CS‐FC modules perform bottleneck‐ and cross‐scale‐guided feature recalibration to suppress noise leakage in skip pathways and enhance decoder reliability. Experimental results on Common Voice and LibriSpeech datasets demonstrate that the proposed DPH‐TFA‐MSFENet achieves superior perceptual evaluation of speech quality, short‐time objective intelligibility, and signal‐to‐distortion ratio performance, particularly under low‐SNR conditions, while maintaining computational efficiency.
No takes yet. Share an insight, caveat, or question.
AreefaBegam et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: