This paper presents a novel framework for automatic speech-driven gesture generation, applicable to human-agent interaction including both virtual agents and robots. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven gesture generation by incorporating representation learning. Our model takes speech as input and produces gestures as output, in the form of a sequence of 3D coordinates.
No takes yet. Share an insight, caveat, or question.
Kucherenko et al. (2019) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: