We present a method that learns to integrate temporal information, from a dynamics model, with ambiguous visual information, from a learned model, in the context of interacting agents. Our method is based on a-structured variational recurrent neural network (Graph-VRNN), which is end-to-end to infer the current state of the (partially observed), as well as to forecast future states. We show that our method various baselines on two sports datasets, one based on real trajectories, and one generated by a soccer game engine.
No takes yet. Share an insight, caveat, or question.
Sun et al. (2019) studied this question.