PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 20260 citationsOpen Access

Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering

View Full Paper
PRPradeep RangappaRVRoberto Andres Carofilis VascoJPJeena Prakash

Key Points

  • The aim is to enhance the adaptation of ASR models in resource-constrained environments by selecting high-quality training data.
  • Developed a data selection pipeline using pseudo-labels from Whisper and Zipformer models
  • Integrated multiple filtering strategies: WER prediction, NER, and CER analysis
  • Evaluated performance against 7500 hours of baseline data followed by a comparative analysis with a CER-based approach
  • Achieved 12.3% WER on 7500 hours of pseudo-labeled call center data
  • Reduced dataset from 7500 hours to 100 hours, maintaining high performance with 1.4% WER
  • Observed similar trends on the Fisher English dataset

Abstract

Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies-including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis-to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rangappa et al. (2025) studied this question.

synapsesocial.com/papers/69be356f6e48c4981c673bc0https://doi.org/10.5167/uzh-293015
Ask AI
Helpful
Bookmark
Share
View Full Paper