Speech-driven lip synchronization is an important technique for talking-face video generation, with broad application potential in virtual humans, video dubbing, digital media, and human–computer interaction. However, existing methods still face challenges in achieving reliable lip synchronization while maintaining stable identity preservation, high visual fidelity, and efficient inference, especially in Chinese-language scenarios where related research remains relatively limited. To address these issues, we propose FastTalk, a speech-driven lip synchronization method for Chinese-language scenarios. The proposed framework performs latent-space restoration for efficient video synthesis, uses a fixed-mask strategy to suppress shortcut visual cues and strengthen audio-driven lip-shape prediction, and adopts a two-stage training scheme to reduce the gap between training and inference. This design improves generation stability while preserving efficiency. Experimental results show that FastTalk achieves competitive lip synchronization performance while improving visual quality and identity preservation. These results indicate that FastTalk provides an effective solution for Chinese speech-driven lip synchronization video generation.
Liu et al. (Fri,) studied this question.