This algorithm improves lip dynamic authenticity in talking head synthesis, suggesting that 3D representations are key.
Audio-driven talking head synthesis aims to generate lifelike facial animations synchronized with audio. Current approaches primarily focus on lip motion information in 2D visual space for lip-audio synchronization and expressive lip dynamic, often neglecting 3D geometric motion representations of the lips that can more accurately capture lip movements in real-world scenarios. This oversight can result in suboptimal lip dynamic authenticity. In this work, we introduce a novel 3D Temporal Representation Learning (3D-TRL) algorithm that models 3D lip temporal information as latent representations and utilizes these representations as additional supervision to enhance dynamic authenticity. To achieve this, we leverage the geometric mesh constructed from the 3D Morphable Model (3DMM) as our 3D information of the lip and explore two self-supervised strategies to learn temporal representation in 3D geometric space. First, we propose a Reconstruction-oriented 3D-TRL algorithm that reconstructs the input to obtain motion tokens in hidden space, encapsulating content while capturing richer contextual representations of the sequence. Second, we develop a Contrastive-based 3D-TRL algorithm that utilizes contrastive learning to extract hidden 3D motion representations. This algorithm employs data augmentation strategies appropriate specifically for the 3D temporal sequences of the lips. Extensive experiments demonstrate that our approach, as a versatile and adaptable supervisory, can be integrated into various state-of-the-art network frameworks, leading to substantial enhancements in lip dynamic authenticity.
No takes yet. Share an insight, caveat, or question.
Li et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: