Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 11, 2025Open Access

TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

View Full Paper
Ask AI
Bookmark
Share

Authors

SCShunian ChenHHHailin HuangYLYexin Liu

Discussion

Loading...

Member takes

Overview

Analysis shows TalkVid improves generalization in audio-driven talking head synthesis, indicating better performance across ethnicity and language demographics.

Key Points

  • A model trained on TalkVid exhibits superior generalization compared to previous datasets, enhancing performance.
  • TalkVid comprises 1244 hours of video from 7729 unique speakers, ensuring diverse representation in the dataset.
  • The dataset was curated through a multi-stage pipeline, focusing on quality and reliability validated by human judgments.
  • TalkVid-Bench includes 500 clips balanced across key demographic and linguistic axes, highlighting performance disparities.

Cite This Study

Chen et al. (2025) studied this question.

synapsesocial.com/papers/68e9b1b5ba7d64b6fc131df4https://doi.org/10.48550/arxiv.2508.13618
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation2025
  2. 2DIVA-3D: a diverse 3D talking head dataset from in-the-wild videos2026
  3. 3MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset2024
  4. 4MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset2024 · 14 citations
  5. 5GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation2025