PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 13, 20041,348 citations

Learning methods for generic object recognition with invariance to pose and lighting

View Full Paper
YLYann LeCunFHFu Jie HuangLBLéon Bottou

Key Points

  • The aim is to evaluate various learning methods for their capacity to recognize objects while maintaining invariance to pose and lighting.
  • Utilized a dataset of 194,400 images of 50 uniform-colored toys under varying conditions.
  • Tested nearest neighbor methods, support vector machines, and convolutional networks on grayscale images.
  • Comparative error rates calculated for unseen instances in controlled and cluttered environments.
  • SVM achieved approximately 13% error on unseen instances with uniform backgrounds.
  • Convolutional networks reported a 7% error rate on uniform backgrounds.
  • In cluttered scenes, SVM was impractical while convolutional nets performed with a 16/7% error rate.

Abstract

We assess the applicability of several popular learning methods for the problem of recognizing generic visual categories with invariance to pose, lighting, and surrounding clutter. A large dataset comprising stereo image pairs of 50 uniform-colored toys under 36 azimuths, 9 elevations, and 6 lighting conditions was collected (for a total of 194,400 individual images). The objects were 10 instances of 5 generic categories: four-legged animals, human figures, airplanes, trucks, and cars. Five instances of each category were used for training, and the other five for testing. Low-resolution grayscale images of the objects with various amounts of variability and surrounding clutter were used for training and testing. Nearest neighbor methods, support vector machines, and convolutional networks, operating on raw pixels or on PCA-derived features were tested. Test error rates for unseen object instances placed on uniform backgrounds were around 13% for SVM and 7% for convolutional nets. On a segmentation/recognition task with highly cluttered images, SVM proved impractical, while convolutional nets yielded 16/7% error. A real-time version of the system was implemented that can detect and classify objects in natural scenes at around 10 frames per second.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

LeCun et al. (2004) studied this question.

synapsesocial.com/papers/6a0712745589773960843536https://doi.org/10.1109/cvpr.2004.1315150
Ask AI
Helpful
Bookmark
Share
View Full Paper