PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 1, 20240 citationsOpen Access

Few-Shot Keyword Spotting from Mixed Speech

View Full Paper
JYJunming YuanYSYing ShiLLLantian Li

Key Points

Key points are not available for this paper at this time.

Abstract

Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples. A commonly used approach is the pre-training and fine-tuning framework. While effective in clean conditions, this approach struggles with mixed keyword spotting – simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications. Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario. In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech. Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase. Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yuan et al. (2024) studied this question.

synapsesocial.com/papers/68e59e8eb6db6435875389b2https://doi.org/10.21437/interspeech.2024-2296
Ask AI
Helpful
Bookmark
Share
View Full Paper