Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 2, 2025Open Access

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

View Full Paper
Ask AI
Bookmark
Share

Authors

YZYouliang ZhangZLZhaoyang LiDWDaQuan Wang

Discussion

Loading...

Member takes

Overview

Dataset enables high-quality generation of dyadic interactive virtual humans, suggesting new research avenues.

Key Points

  • SpeakerVid-5M provides over 8,743 hours of video clips, which facilitates the study of interactive virtual humans.
  • The dataset contains more than 5.2 million video clips structured by interaction type and data quality.
  • It includes a high-quality subset for supervised fine-tuning, enhancing the performance of virtual human systems.
  • An autoregressive-based video chat baseline and benchmarks have been established for future research and comparisons.

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68de5da783cbc991d0a20cbahttps://doi.org/10.48550/arxiv.2507.09862
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis2025
  2. 2OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation2024 · 2 citations
  3. 3Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation2024
  4. 4Voxblink: A Large Scale Speaker Verification Dataset on Camera2024 · 19 citations
  5. 5DIVA-3D: a diverse 3D talking head dataset from in-the-wild videos2026