PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026The Journal of the Acoustical Society of America0 citations

Pop-out voice generation via acoustic feature control in text-to-speech synthesis

View Full Paper
RURyosuke UesugiMeijo UniversityHBHideki BannoMeijo UniversityKAKensaku AsahiMeijo University

Key Points

  • This research aims to enhance pop-out voice generation in synthetic speech by manipulating acoustic features.
  • Utilized JETS, an end-to-end text-to-speech model, to incorporate a dynamic feature (DF) predictor alongside existing predictors.
  • Conducted preliminary evaluations comparing synthetic samples with original recordings to assess quality and DF levels.
  • Explored an approach emphasizing DF output by adjusting emphasis weighting to enhance pop-out qualities.
  • The integration of the DF predictor did not degrade synthetic speech quality, with mean DF values similar to original recordings.
  • Emphasizing the DF predictor led to higher DF levels in generated samples, although some degradation in speech quality was observed.

Abstract

Speech that stands out from background noise is called a “pop-out” voice. Several acoustic features are known to contribute to the degree of pop-out, including fundamental frequency, power in the frequency band above 1 kHz, and dynamic feature (DF). Our research focuses on controlling the degree of pop-out in Text-to-Speech (TTS) by modifying relevant acoustic features, using JETS, which is an end-to-end TTS model, as the base. We first attempted to incorporate DF into the model by adding a “DF predictor” alongside the existing pitch, energy, and duration predictors in the variance adaptor of JETS. Preliminary evaluations of the model without explicit DF control showed that adding the DF predictor did not degrade the subjective quality of the synthetic samples, and the mean DF values of the synthetic samples were comparable to those of the original recordings. To produce more pop-out-like samples, we introduced an approach that emphasizes the output of the DF predictor. This method generated samples with higher DF when the emphasis weighting was set to a large value. However, some degradation was observed in the generated speech. We are currently working to address this issue.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Uesugi et al. (2025) studied this question.

synapsesocial.com/papers/6a05677ca550a87e60a1f8dbhttps://doi.org/10.1121/10.0041580
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Perceptual and acoustic characteristics of pop-out voice2026
  2. 2Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis2025
  3. 3Speaker-dependent differences in effective acoustic features for deepfake speech detection using identical English words2025
  4. 4Characterizing Sustained Phonation in Text-To-Speech Models2026
  5. 5Beyond graphemes and phonemes: continuous phonological features in neural text-to-speech synthesis2024