Key points are not available for this paper at this time.
3D facial reconstruction and animation has advanced significantly in the past decades. However, existing methods based on single-modal input struggle with specific facial part control and require post-processing for natural rendering. Recent approaches integrate audio and 2D estimation results to enhance facial animations, which improves the naturalness of head pose, eye and mouth animations. In this paper, We present a multi-modal 3D facial animation framework that uses video and audio input simultaneously. We experiment on different settings of motion control to produce natural and precise facial animation through user studies.
Cao et al. (Thu,) studied this question.