PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 17, 2025ACM Transactions on Asian and Low-Resource Language Information Processing1 citations

Enhanced Prosody Modeling and Character Voice Controlling for Audiobook Speech Synthesis

View Full Paper
NWNing-Qian WuZLZhen-Hua Ling

Key Points

  • Enhanced prosody modeling improves the expressiveness of synthesized audiobook speech, allowing for more human-like performance.
  • Incorporating generative adversarial networks for phoneme-level prosody code prediction significantly upgrades synthesis quality.
  • The character voice encoder produces distinct voices for different characters, enhancing character dialogue delivery.
  • Experimental validation indicates that the new approach maintains the naturalness of synthesized speech while providing rich expressiveness.

Abstract

Conventional speech synthesis techniques have made significant strides towards achieving human-like performance. However, the domain of audiobook speech synthesis still presents notable challenges. On one hand, the speech in audiobooks exhibits rich prosodic expressiveness, posing substantial difficulties in prosody modeling. On the other hand, the reader of audiobooks uses different voices to perform dialogues of different characters, which has been inadequately explored in existing speech synthesis methods. To address the first challenge, we integrate discourse-scale prosody modeling into the conventional autoencoder-based framework and introduce generative adversarial networks (GANs) for phoneme-level prosody code prediction. Regarding the second challenge, we further explore a character voice encoder based on the pretrained speaker verification model, integrating it into our proposed method. Experimental results validate that the proposed method enhances the prosodic expressiveness of synthesized audiobook speech. Moreover, it demonstrates the capacity to produce distinctive voices for different audiobook characters without compromising the naturalness of the synthesized speech.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68a36f8a0a429f797333261dhttps://doi.org/10.1145/3749644
Ask AI
Helpful
Bookmark
Share
View Full Paper