PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 25, 20240 citationsOpen Access

MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization

View Full Paper
AFAdriana Fernandez-LopezHCHonglie ChenPMPingchuan Ma

Key Points

  • Sparse mask optimization enables training multimodal speech recognition from scratch, achieving 21.1% visual and 0.9% audio-visual word error rates on the LRS3 benchmark.
  • Benchmark on the LRS3 dataset applies sparse regularization early in training to maintain gradient flow, before transitioning to dense models or maintaining sparse masks.
  • This regularization approach reduces training time by over twofold, while overcoming vanishing gradients to eliminate reliance on computationally costly pre-trained models.

Abstract

Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facilitates the training of visual and audio-visual speech recognition models (VSR and AVSR) from scratch. This approach, abbreviated as MSRS (Multimodal Speech Recognition from Scratch), introduces a sparse regularization that rapidly learns sparse structures within the dense model at the very beginning of training, which receives healthier gradient flow than the dense equivalent. Once the sparse mask stabilizes, our method allows transitioning to a dense model or keeping a sparse model by updating non-zero values. MSRS achieves competitive results in VSR and AVSR with 21. 1% and 0. 9% WER on the LRS3 benchmark, while reducing training time by at least 2x. We explore other sparse approaches and show that only MSRS enables training from scratch by implicitly masking the weights affected by vanishing gradients.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fernandez-Lopez et al. (2024) studied this question.

synapsesocial.com/papers/68e636c5b6db6435875c8bc8https://doi.org/10.48550/arxiv.2406.17614
Ask AI
Helpful
Bookmark
Share
View Full Paper