PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 20250 citationsOpen Access

AvatarSync: Rethinking Talking-Head Animation through Autoregressive Perspective

View Full Paper
YDYuchen DengXWXianliang WuHZHai-Tao Zheng

Key Points

  • AvatarSync improves visual fidelity and computational efficiency in talking-head animations.
  • The method employs a two-stage generation strategy to enhance temporal coherence and the representation of phonemes.
  • Extensive experiments show AvatarSync outperforms previous methods in reducing inter-frame flicker and identity drift.
  • Optimizing the inference pipeline significantly reduces latency while maintaining visual fidelity.

Abstract

Existing talking-head animation approaches based on Generative Adversarial Networks (GANs) or diffusion models often suffer from inter-frame flicker, identity drift, and slow inference. These limitations inherent to their video generation pipelines restrict their suitability for applications. To address this, we introduce AvatarSync, an autoregressive framework on phoneme representations that generates realistic and controllable talking-head animations from a single reference image, driven directly text or audio input. In addition, AvatarSync adopts a two-stage generation strategy, decoupling semantic modeling from visual dynamics, which is a deliberate "Divide and Conquer" design. The first stage, Facial Keyframe Generation (FKG), focuses on phoneme-level semantic representation by leveraging the many-to-one mapping from text or audio to phonemes. A Phoneme-to-Visual Mapping is constructed to anchor abstract phonemes to character-level units. Combined with a customized Text-Frame Causal Attention Mask, the keyframes are generated. The second stage, inter-frame interpolation, emphasizes temporal coherence and visual smoothness. We introduce a timestamp-aware adaptive strategy based on a selective state space model, enabling efficient bidirectional context reasoning. To support deployment, we optimize the inference pipeline to reduce latency without compromising visual fidelity. Extensive experiments show that AvatarSync outperforms existing talking-head animation methods in visual fidelity, temporal consistency, and computational efficiency, providing a scalable and controllable solution.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Deng et al. (2025) studied this question.

synapsesocial.com/papers/68dc12cc8a7d58c25ebb0c6bhttps://doi.org/10.48550/arxiv.2509.12052
Ask AI
Helpful
Bookmark
Share
View Full Paper