PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

View Full Paper
XHXiluo HeAPAlexander PolokJVJesús Villalba

Key Points

  • The proposed method reduces inference costs in multi-talker automatic speech recognition while maintaining performance.
  • By converting speaker-specific activity outputs into streams, we avoid scaling costs with the number of speakers.
  • New heuristics were designed to ensure compatibility with existing systems while preserving conversational continuity.
  • Results demonstrate that this approach, compatible with Diarization-Conditioned Whisper, reduces runtimes on the AMI and ICSI datasets.

Abstract

An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that naïvely merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

He et al. (2025) studied this question.

synapsesocial.com/papers/68e865117ef2f04ca37e4d45https://doi.org/10.48550/arxiv.2510.03630
Ask AI
Helpful
Bookmark
Share
View Full Paper