Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
June 30, 2026Visual IntelligenceOpen Access

DIVA-3D: a diverse 3D talking head dataset from in-the-wild videos

View Full Paper
Ask AI
Bookmark
Share

Authors

YWYuhan WuBeijing University of Posts and TelecommunicationsYZYixuan ZhangCentre for Artificial Intelligence and RoboticsQCQing ChangZhejiang University of Science and Technology

Discussion

Loading...

Member takes

Implication

Randomized trial evaluates lip synchronization and facial expressions in a diverse 3D talking head dataset, highlighting significant improvements over existing methods.

Key Points

  • The central aim is to create a diverse dataset for training lifelike 3D talking head models that improve lip synchronization and facial expressions.
  • Developed a semi-automated pipeline to collect audio and corresponding 3D facial FLAME data from public videos.
  • Constructed DIVA-3D, a large-scale audio-visual dataset containing 73 hours of data in Chinese and English.
  • Benchmarked state-of-the-art methods against the new dataset to validate effectiveness.
  • DIVA-3D demonstrates superior performance in generating accurate lip synchronization with an effectiveness rate exceeding prior models.
  • Comprehensive benchmarks indicate that the proposed generative framework significantly enhances facial expressions accuracy.
  • The dataset is the most topically diverse available, encompassing six distinct domains.

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/6a435c8c759b888809a52fddhttps://doi.org/10.1007/s44267-026-00120-6
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A 3-D Audio-Visual Corpus of Affective Communication2010 · 161 citations
  2. 2Synthesizing Obama2017 · 1,097 citations
  3. 3SPECTRE: Visual Speech-Informed Perceptual 3D Facial Expression Reconstruction from Videos2023 · 45 citations
  4. 4Audio-driven facial animation by joint end-to-end learning of pose and emotion2017 · 449 citations
  5. 5VoxCeleb: A Large-Scale Speaker Identification Dataset2017 · 2,184 citations