PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 25, 2020148 citationsOpen Access

Spot the Conversation: Speaker Diarisation in the Wild

JCJoon Son ChungJHJaesung HuhANArsha Nagrani

Key Points

Key points are not available for this paper at this time.

Abstract

The goal of this paper is speaker diarisation of videos collected ‘in the wild’. make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate videos with diarisation labels. Finally, we use this pipeline to create a large-scale diarisation dataset called VoxConverse, collected from ‘in the wild’ videos, which we will release publicly to the research community. Our dataset consists of overlapping speech, a large and diverse speaker pool, and challenging background conditions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chung et al. (2020) studied this question.

synapsesocial.com/papers/6a16cdbdf3be5e880d6b8f5ahttps://doi.org/10.21437/interspeech.2020-2337
Ask AI
Helpful
Bookmark
Share
View Full Paper