PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 20240 citationsOpen Access

Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

View Full Paper
ZDZongyang DuJLJunchen LuKZKun Zhou

Key Points

  • Expressive voice conversion successfully transfers speaker identity and emotional style without separate vocoders by using a conditional denoising diffusion probabilistic model.
  • The end-to-end framework incorporates self-supervised speech units for content alongside features from speech emotion recognition and speaker verification systems.
  • Objective and subjective evaluations demonstrate robust emotion prosody modeling, indicating a viable method for expressive speech conversion across arbitrary voices.

Abstract

Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Du et al. (2024) studied this question.

synapsesocial.com/papers/68e6beabb6db64358763efcahttps://doi.org/10.48550/arxiv.2405.01730
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Converting Anyone's Voice: End-to-End Expressive Voice Conversion with A Conditional Diffusion Model2024 · 4 citations
  2. 2DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion2024 · 34 citations
  3. 3DiffVC+: Improving Diffusion-based Voice Conversion for Speaker Anonymization2024 · 4 citations
  4. 4ESVC: Combining Adaptive Style Fusion and Multi-Level Feature Disentanglement for Expressive Singing Voice Conversion2024 · 2 citations
  5. 5Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework2025