PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 202420 citationsOpen Access

ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

View Full Paper
HTHaobin TangXZXulong ZhangNCNing Cheng

Key Points

Key points are not available for this paper at this time.

Abstract

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (SED) and Speech Emotion Recognition (SER) to model emotions at different levels. Specifically, our proposed approach integrates the utterance-level emotion embedding extracted by SER with fine-grained frame-level emotion embedding obtained from SED. These embeddings are used to condition the reverse process of the denoising diffusion probabilistic model (DDPM). Additionally, we employ cross-domain SED to accurately predict soft labels, addressing the challenge of a scarcity of fine-grained emotion-annotated datasets for supervising emotional TTS training.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tang et al. (2024) studied this question.

synapsesocial.com/papers/68e7376bb6db6435876b0ea1https://doi.org/10.1109/icassp48485.2024.10446467
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis2024 · 6 citations
  2. 2Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition2024
  3. 3DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech2025
  4. 4TEA-VITS: emotion voice synthesis based on temporal emotion analysis2024 · 1 citations
  5. 5Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement2025