Model-based human pose estimation is currently approached through two paradigms. Optimization-based methods fit a parametric body model to2D observations in an iterative manner, leading to accurate image-model, but are often slow and sensitive to the initialization. In, regression-based methods, that use a deep network to directly the model parameters from pixels, tend to provide reasonable, but not accurate, results while requiring huge amounts of supervision. In this, instead of investigating which approach is better, our key insight is the two paradigms can form a strong collaboration. A reasonable, directly estimate from the network can initialize the iterative optimization the fitting faster and more accurate. Similarly, a pixel accurate fit iterative optimization can act as strong supervision for the network. This the core of our proposed approach SPIN (SMPL oPtimization IN the loop). The network initializes an iterative optimization routine that fits the body to 2D joints within the training loop, and the fitted estimate is used to supervise the network. Our approach is self-improving by, since better network estimates can lead the optimization to better, while more accurate optimization fits provide better supervision for network. We demonstrate the effectiveness of our approach in different, where 3D ground truth is scarce, or not available, and we outperform the state-of-the-art model-based pose estimation by significant margins. The project website with videos, results, code can be found at https://seas.upenn.edu/~nkolot/projects/spin.
No takes yet. Share an insight, caveat, or question.
Kolotouros et al. (2019) studied this question.