Virtual conversational agents are supposed to combine speech with non-verbal modalities for intelligible and believable utterances. However, the automatic synthesis of co-verbal gestures is still struggling with several problems like naturalness in procedurally generated animations, flexibility in pre-defined movements, and synchronization with speech. In this paper we focus on generating complex multimodal utterances including gesture and speech from XML-based descriptions of their overt form. We describe a coordination model that reproduces coarticulation and transition effects in both modalities. In particular, an efficient kinematic approach to creating gesture animations from shape specifications is presented, which provides fine adaptation to temporal constraints that are imposed by cross-modal synchrony.
No takes yet. Share an insight, caveat, or question.
Kopp et al. (2003) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: