PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 2018IEEE Transactions on Pattern Analysis and Machine Intelligence195 citationsOpen Access

Fine-Tuning CNN Image Retrieval with No Human Annotation

FRFilip RadenovićGTGiorgos ToliasOCOndřej Chum

Key Points

  • The aim is to enhance image retrieval performance by fine-tuning CNNs using unannotated images and 3D model guidance.
  • Fine-tuned CNNs on unordered images with no human annotation.
  • Utilized reconstructed 3D models to select training data with hard-positive and hard-negative examples.
  • Implemented a novel Generalized-Mean pooling layer to improve performance.
  • Achieved state-of-the-art performance on benchmarks such as Oxford Buildings, Paris, and Holidays datasets.
  • Descriptor whitening from CNN training data outperformed PCA whitening.
  • Enhanced retrieval performance noted through the use of hard-positive and hard-negative selection.

Abstract

Image descriptors based on activations of Convolutional Neural Networks (CNNs) have become dominant in image retrieval due to their discriminative power, compactness of representation, and search efficiency. Training of CNNs, either from scratch or fine-tuning, requires a large amount of annotated data, where a high quality of annotation is often crucial. In this work, we propose to fine-tune CNNs for image retrieval on a large collection of unordered images in a fully automated manner. Reconstructed 3D models obtained by the state-of-the-art retrieval and structure-from-motion methods guide the selection of the training data. We show that both hard-positive and hard-negative examples, selected by exploiting the geometry and the camera positions available from the 3D models, enhance the performance of particular-object retrieval. CNN descriptor whitening discriminatively learned from the same training data outperforms commonly used PCA whitening. We propose a novel trainable Generalized-Mean (GeM) pooling layer that generalizes max and average pooling and show that it boosts retrieval performance. Applying the proposed method to the VGG network achieves state-of-the-art performance on the standard benchmarks: Oxford Buildings, Paris, and Holidays datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Radenović et al. (2018) studied this question.

synapsesocial.com/papers/6a093977b7dd28a06e1611dehttps://doi.org/10.1109/tpami.2018.2846566
Ask AI
Helpful
Bookmark
Share
View Full Paper