This work addresses multi-class segmentation of indoor scenes with RGB-D. While this area of research has gained much attention recently, most still rely on hand-crafted features. In contrast, we apply a multiscale network to learn features directly from the images and the depth. We obtain state-of-the-art on the NYU-v2 depth dataset with an of 64.5%. We illustrate the labeling of indoor scenes in videos that could be processed in real-time using appropriate hardware such an FPGA.
No takes yet. Share an insight, caveat, or question.
Couprie et al. (2013) studied this question.