PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 19, 20250 citationsOpen Access

AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning

View Full Paper
DZDaning ZhangYWYing Wei

Key Points

  • AVFSNet provides a state-of-the-art solution for audio-visual speech separation, especially in scenarios with unknown speaker quantities.
  • The model integrates multi-scale encoding and parallel architecture, enhancing its adaptability to environmental noise.
  • Comprehensive evaluations across multiple datasets reveal that AVFSNet excels in both separation tasks and speaker counting capabilities.
  • Existing methods struggle with unknown speaker count, but AVFSNet significantly improves generalization in real-world acoustic settings.

Abstract

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge of speaker counts in mixtures. The limited research addressing unknown speaker quantity scenarios exhibits significantly constrained generalization capabilities in real acoustic environments. To overcome these challenges, this paper proposes AVFSNet -- an audio-visual speech separation model integrating multi-scale encoding and parallel architecture -- jointly optimized for speaker counting and multi-speaker separation tasks. The model independently separates each speaker in parallel while enhancing environmental noise adaptability through visual information integration. Comprehensive experimental evaluations demonstrate that AVFSNet achieves state-of-the-art results across multiple evaluation metrics and delivers outstanding performance on diverse datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68f4b10d3d9d770bbc697091https://doi.org/10.48550/arxiv.2507.12972
Ask AI
Helpful
Bookmark
Share
View Full Paper