PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20242 citationsOpen Access

On Real-Time Multi-Stage Speech Enhancement Systems

View Full Paper
LMLingjun MengJCJozef ColdenhoffPKPaul Kendrick

Key Points

Key points are not available for this paper at this time.

Abstract

Recently, multi-stage systems have stood out among deep learning-based speech enhancement methods. However, these systems are always high in complexity, requiring millions of parameters and powerful computational resources, which limits their application for real-time processing in low-power devices. Besides, the contribution of various influencing factors to the success of multi-stage systems remains unclear, which presents challenges to reduce the size of these systems. In this paper, we extensively investigate a lightweight two-stage network with only 560k total parameters. It consists of a Mel-scale magnitude masking model in the first stage and a complex spectrum mapping model in the second stage. We first provide a consolidated view of the roles of gain power factor, post-filter, and training labels for the Mel-scale masking model. Then, we explore several training schemes for the two-stage network and provide some insights into the superiority of the two-stage network. We show that the proposed two-stage network trained by an optimal scheme achieves a performance similar to a four times larger open source model DeepFilterNet2 1.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Meng et al. (2024) studied this question.

synapsesocial.com/papers/68e7398bb6db6435876b2c9chttps://doi.org/10.1109/icassp48485.2024.10447228
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement2025
  2. 2Dual‐Path Higher‐Order Information Interaction With Time‐Frequency Attention and Multi‐Scale Feature Extraction Based Nested U‐Net for Speech Enhancement2026
  3. 3A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet22024
  4. 4Light-Weight Causal Speech Enhancement Using Time-Varying Multi-Resolution Filtering2024
  5. 5Stacked Multiscale Densely Connected Temporal Convolutional Attention Network for Multi-Objective Speech Enhancement in an Airborne Environment2024 · 1 citations