PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 27, 201690 citationsOpen Access

Multi-modal Auto-Encoders as Joint Estimators for Robotics Scene Understanding

CCCésar CadenaADAnthony DickIRIan Reid

Key Points

Key points are not available for this paper at this time.

Abstract

We explore the capabilities of Auto-Encoders to fuse the information available from cameras and depth sensors, and to reconstruct missing data, for scene understanding tasks. In particular we consider three input modalities: RGB images; depth images; and semantic label information. We seek to generate complete scene segmentations and depth maps, given images and partial and/or noisy depth and semantic data. We formulate this objective of reconstructing one or more types of scene data using a Multi-modal stacked Auto-Encoder. We show that suitably designed Multi-modal Auto-Encoders can solve the depth estimation and the semantic segmentation problems simultaneously, in the partial or even complete absence of some of the input modalities. We demonstrate our method using the outdoor dataset KITTI that includes LIDAR and stereo cameras. Our results show that as a means to estimate depth from a single image, our method is comparable to the state-of-the-art, and can run in real time (i.e., less than 40ms per frame). But we also show that our method has a significant advantage over other methods in that it can seamlessly use additional data that may be available, such as a sparse point-cloud and/or incomplete coarse semantic labels.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cadena et al. (2016) studied this question.

synapsesocial.com/papers/6a1ae18b739ab56a9086158bhttps://doi.org/10.15607/rss.2016.xii.041
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Caffe: Convolutional Architecture for Fast Feature Embedding2014 · 4,306 citations
  2. 2Deep Learning of Representations for Unsupervised and Transfer Learning.2011 · 893 citations
  3. 3Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion2010 · 4,136 citations
  4. 4Pulling Things out of Perspective2014 · 502 citations
  5. 5Depth Map Prediction from a Single Image using a Multi-Scale Deep Network2014 · 2,267 citations