PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 26, 202439 citationsOpen Access

AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset

ZCZhixi CaiSGShreya GhoshAAAman Pankaj Adatia

Key Points

Key points are not available for this paper at this time.

Abstract

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small segments of audio-visual manipulations embedded in real videos. In this research, we emulate the process of such content generation and propose the AV-Deepfake1M dataset. The dataset contains content-driven (i) video manipulations, (ii) audio manipulations, and (iii) audio-visual manipulations for more than 2K subjects resulting in a total of more than 1M videos. The paper provides a thorough description of the proposed data generation pipeline accompanied by a rigorous analysis of the quality of the generated data. The comprehensive benchmark of the proposed dataset utilizing state-of-the-art deepfake detection and localization methods indicates a significant drop in performance compared to previous datasets. The proposed dataset will play a vital role in building the next-generation deepfake localization methods. The dataset and associated code are available at https://github.com/ControlNet/AV-Deepfake1M.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cai et al. (2024) studied this question.

synapsesocial.com/papers/6a1f9721f184cd72a625ec3fhttps://doi.org/10.1145/3664647.3680795
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Applied Soft Computing2021 · 546 citations
  2. 2Converting video formats with FFmpeg2006 · 282 citations
  3. 3Xception: Deep Learning with Depthwise Separable Convolutions2017 · 19,482 citations
  4. 4Powerset multi-class cross entropy loss for neural speaker diarization2023 · 115 citations
  5. 5Detecting Deepfakes with Self-Blended Images2022 · 444 citations