High quality speech-to-lips conversion, investigated in this work, ren-ders realistic lips movement (video) consistent with input speech (audio) without knowing its linguistic content. Instead of memoryless frame-based conversion, we adopt maximum likelihood estimation of the vi-sual parameter trajectories using an audio-visual joint Gaussian Mixture Model (GMM). We propose a minimum converted trajectory error ap-proach (MCTE) to further refine the converted visual parameters. First, we reduce the conversion error by training the joint audio-visual GMM with weighted audio and visual likelihood. Then MCTE uses the gen-eralized probabilistic descent algorithm to minimize a conversion error of the visual parameter trajectories defined on the optimal Gaussian ker-nel sequence according to the input speech. We demonstrate the effec-tiveness of the proposed methods using the LIPS 2009 Visual Speech Synthesis Challenge dataset, without knowing the linguistic (phonetic) content of the input speech. Index Terms: visual speech synthesis, speech-to-lips conversion, mini-mum conversion error, minimum generation error
No takes yet. Share an insight, caveat, or question.
Zhuang et al. (2010) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: