PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 1, 202411 citations

Spoofed Speech Detection with a Focus on Speaker Embedding

View Full Paper
HTHoan My TranDGDavid GuennecPMPhilippe Martin

Key Points

Key points are not available for this paper at this time.

Abstract

Self-Supervised Learning (SSL) models excel as feature extractors in downstream speech tasks, including the increasingly crucial area of spoof speech detection due to the rise of audio deepfakes using Text-To-Speech (TTS) and Voice Conversion (VC) technologies. To address this issue, we propose a novel approach that relies on speaker embedding using a finetuned WavLM model with layer-wise attentive statistics pooling combined to a supervised contrastive learning and cross-entropy loss. Evaluation on Logical Access (LA) and DeepFake (DF) tasks on ASVspoof 2019 and 2021 highlights its potential in detecting audio deepfakes, with the contrastive loss producing more stable results among test sets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tran et al. (2024) studied this question.

synapsesocial.com/papers/68e59e8eb6db64358753866chttps://doi.org/10.21437/interspeech.2024-481
Ask AI
Helpful
Bookmark
Share
View Full Paper