PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 2025Electronics0 citationsOpen Access

FFMamba: Feature Fusion State Space Model Based on Sound Event Localization and Detection

View Full Paper
YLYibo LiDGDongyuan GeJXJianqiang Xu

Key Points

  • The FFMamba model effectively captures long-range temporal dependencies and integrates multi-scale audio features.
  • Comparative experiments showed significant performance improvements on DCASE2021 and DCASE2022 datasets.
  • Ablation studies confirmed the essential contributions of the MSFVSS and WTED modules in enhancing model performance.
  • This novel state space model combines local spatial details with long-range dependency capture, outperforming traditional CNN and Transformer approaches.

Abstract

Previous studies on Sound Event Localization and Detection (SELD) have primarily focused on CNN- and Transformer-based designs. While CNNs possess local receptive fields, making it difficult to capture global dependencies over long sequences, Transformers excel at modeling long-range dependencies but have limited sensitivity to local time–frequency features. Recently, the VMamba architecture, built upon the Visual State Space (VSS) model, has shown great promise in handling long sequences, yet it remains limited in modeling local spatial details. To address this issue, we propose a novel state space model with an attention-enhanced feature fusion mechanism, termed FFMamba, which balances both local spatial modeling and long-range dependency capture. At a fine-grained level, we design two key modules: the Multi-Scale Fusion Visual State Space (MSFVSS) module and the Wavelet Transform-Enhanced Downsampling (WTED) module. Specifically, the MSFVSS module integrates a Multi-Scale Fusion (MSF) component into the VSS framework, enhancing its ability to capture both long-range temporal dependencies and detailed local spatial information. Meanwhile, the WTED module employs a dual-branch design to fuse spatial and frequency domain features, improving the richness of feature representations. Comparative experiments were conducted on the DCASE2021 Task 3 and DCASE2022 Task 3 datasets. The results demonstrate that the proposed FFMamba model outperforms recent approaches in capturing long-range temporal dependencies and effectively integrating multi-scale audio features. In addition, ablation studies confirmed the effectiveness of the MSFVSS and WTED modules.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68dc1e308a7d58c25ebb130ehttps://doi.org/10.3390/electronics14193874
Ask AI
Helpful
Bookmark
Share
View Full Paper