PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20241 citationsOpen Access

Cross-Modal Alignment for End-to-End Spoken Language Understanding Based on Momentum Contrastive Learning

View Full Paper
BZBeida ZhengMAMijit AblimitAHAskar Hamdulla

Key Points

Key points are not available for this paper at this time.

Abstract

The end-to-end spoken language understanding system extracts the semantic intent directly from an input speech. It effectively avoids problems such as semantic drift in traditional cascade models. However, the lack of semantically labeled speech data makes the model training process diffi-cult. Several recent multi-modal research perspectives have demonstrated that aligning speech and text embeddings based on space distance can improve the model's performance. In this study, inspired by the work related to contrastive learning, a speech and text aligning method using momentum contrast learning is proposed, and a momentum distillation method is also used in the model to learn from imperfectly matched speech and text data. The proposed method has improved intent detection accuracy by 2.14% and 5.98% on Fluent Speech Command and SmartLights datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zheng et al. (2024) studied this question.

synapsesocial.com/papers/68e7397eb6db6435876b29e9https://doi.org/10.1109/icassp48485.2024.10448143
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Jaco: An Offline Running Privacy-aware Voice Assistant2022 · 13 citations
  2. 2Snips Voice Platform: an embedded Spoken Language Understanding system\n for private-by-design voice interfaces2018 · 355 citations
  3. 3Improving End-to-End Speech-to-Intent Classification with Reptile2020 · 23 citations
  4. 4Finstreder: Simple and fast Spoken Language Understanding with Finite State Transducers using modern Speech-to-Text models2022 · 3 citations
  5. 5Adam: A Method for Stochastic Optimization2014 · 84,704 citations