PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 8, 2024Neural Processing Letters18 citationsOpen Access

Multi-view Self-supervised Learning and Multi-scale Feature Fusion for Automatic Speech Recognition

View Full Paper
JZJingyu ZhaoRLRuwei LiMTMaocun Tian

Key Points

Key points are not available for this paper at this time.

Abstract

Abstract To address the challenges of the poor representation capability and low data utilization rate of end-to-end speech recognition models in deep learning, this study proposes an end-to-end speech recognition model based on multi-scale feature fusion and multi-view self-supervised learning (MM-ASR). It adopts a multi-task learning paradigm for training. The proposed method emphasizes the importance of inter-layer information within shared encoders, aiming to enhance the model’s characterization capability via the multi-scale feature fusion module. Moreover, we apply multi-view self-supervised learning to effectively exploit data information. Our approach is rigorously evaluated on the Aishell-1 dataset and further validated its effectiveness on the English corpus WSJ. The experimental results demonstrate a noteworthy 4. 6 \% % reduction in character error rate, indicating significantly improved speech recognition performance. These findings showcase the effectiveness and potential of our proposed MM-ASR model for end-to-end speech recognition tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2024) studied this question.

synapsesocial.com/papers/68e6b01bb6db6435876316fdhttps://doi.org/10.1007/s11063-024-11614-z
Ask AI
Helpful
Bookmark
Share
View Full Paper