Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 20, 2025Open Access

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

View Full Paper
Ask AI
Bookmark
Share

Authors

KGKe GuZWZhicong WuPBPeng Bai

Discussion

Loading...

Member takes

Overview

PerformSinger demonstrates enhanced audio synthesis in singing voice models by incorporating lip cues, suggesting improved duration prediction accuracy.

Key Points

  • PerformSinger enables high-quality, duration-free singing voice synthesis, integrating visual lip cues.
  • The framework includes advanced components like multimodal encoders and a vocoder, enhancing audio quality.
  • Extensive experiments showcase its state-of-the-art performance in both subjective and objective evaluations.
  • A novel dataset was created with synchronized videos and precise phoneme annotations to support this research.

Cite This Study

Gu et al. (2025) studied this question.

synapsesocial.com/papers/68f6196ee0bbbc94fac36333https://doi.org/10.48550/arxiv.2509.22718
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture2025
  2. 2Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model2024 · 4 citations
  3. 3Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis2024
  4. 4X-Singer: Code-Mixed Singing Voice Synthesis via Cross-Lingual Learning2024 · 1 citations
  5. 5VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation2024