PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20243 citationsOpen Access

Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models

View Full Paper
CSChristopher SimicTBTobias Bocklet

Key Points

Key points are not available for this paper at this time.

Abstract

Automatic speech recognition (ASR) has reached a level of accuracy in recent years, that even outperforms humans in transcribing speech to text. Nevertheless, all current ASR approaches show a certain weakness against ambient noise. To reduce this weakness, audio-visual speech recognition (AVSR) approaches additionally consider visual information from lip movements for transcription. This additional modality increases the computational cost for training models from scratch. We propose an approach, that builds on a pre-trained ASR model and extends it with an adaptive upstream module, that fuses audio and visual information. Since we do not need to train the transformer structure from scratch, our approach requires a fraction of the computational resources compared to traditional AVSR models. Compared to current SOTA systems like AV-HuBERT, our approach achieves an average improvement of 8.3 % in word error rate across different model sizes, noise categories and broad SNR range. The approach allows up to 21 % smaller models and requires only a fraction of the computational resources for training and inference compared to common AVSR approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Simic et al. (2024) studied this question.

synapsesocial.com/papers/68e7387fb6db6435876b161fhttps://doi.org/10.1109/icassp48485.2024.10448047
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LRS3-TED: a large-scale dataset for visual speech recognition2018 · 281 citations
  2. 2Robust Speech Recognition via Large-Scale Weak Supervision2022 · 1,187 citations
  3. 3Perceptual Signal Analysis and Eigenvalue-Weighted Metrics for Musical Audio Quality Assessment2026 · 918 citations
  4. 4SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition2019 · 3,620 citations
  5. 52008 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)2007 · 217 citations