Understanding the 3D world is a fundamental problem in computer vision., learning a good representation of 3D objects is still an open problem to the high dimensionality of the data and many factors of variation. In this work, we investigate the task of single-view 3D object from a learning agent's perspective. We formulate the learning as an interaction between 3D and 2D representations and propose an-decoder network with a novel projection loss defined by the perspective. More importantly, the projection loss enables the unsupervised using 2D observation without explicit 3D supervision. We demonstrate ability of the model in generating 3D volume from a single 2D image with sets of experiments: (1) learning from single-class objects; (2) learning multi-class objects and (3) testing on novel object classes. Results show performance and better generalization ability for 3D object when the projection loss is involved.
No takes yet. Share an insight, caveat, or question.
Yan et al. (2016) studied this question.