PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 24, 20250 citationsOpen Access

EmoCAST: Emotional Talking Portrait via Emotive Text Description

View Full Paper
YJYiguo JiangXCXiaodong CunYZYong Zhang

Key Points

  • EmoCAST achieves state-of-the-art performance in generating realistic talking portrait videos, enhancing emotional expression.
  • The model utilizes a text-guided emotive module along with an emotive audio attention module to improve synthesis accuracy.
  • A novel emotional talking head dataset is constructed, optimizing the framework's ability to capture nuanced features.
  • The proposed training strategies significantly improve lip synchronization and emotional alignment in generated videos.

Abstract

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available datasets are primarily collected in lab settings, further exacerbating these shortcomings. Consequently, these limitations substantially hinder practical applications in real-world scenarios. To address these challenges, we propose EmoCAST, a diffusion-based framework with two key modules for precise text-driven emotional synthesis. In appearance modeling, emotional prompts are integrated through a text-guided decoupled emotive module, enhancing the spatial knowledge to improve emotion comprehension. To improve the relationship between audio and emotion, we introduce an emotive audio attention module to capture the interplay between controlled emotion and driving audio, generating emotion-aware features to guide more precise facial motion synthesis. Additionally, we construct an emotional talking head dataset with comprehensive emotive text descriptions to optimize the framework's performance. Based on the proposed dataset, we propose an emotion-aware sampling training strategy and a progressive functional training strategy that further improve the model's ability to capture nuanced expressive features and achieve accurate lip-synchronization. Overall, EmoCAST achieves state-of-the-art performance in generating realistic, emotionally expressive, and audio-synchronized talking-head videos. Project Page: https://github.com/GVCLab/EmoCAST

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2025) studied this question.

synapsesocial.com/papers/68d6e0fc8b2b6861e4c3f4fehttps://doi.org/10.48550/arxiv.2508.20615
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1EmoVOCA: Speech-Driven Emotional 3D Talking Heads2024 · 1 citations
  2. 2EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head2024
  3. 3EmoFace: Audio-driven Emotional 3D Face Animation2024 · 16 citations
  4. 4EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions2024 · 3 citations
  5. 5Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation2026 · 1 citations