PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 31, 2022IEEE Robotics and Automation Letters27 citations

A Deep Feature Aggregation Network for Accurate Indoor Camera Localization

View Full Paper
TXTao XieKDKun DaiKWKe Wang

Key Points

  • This work aims to enhance indoor camera localization accuracy using low-level feature maps.
  • Developed a Deep Feature Aggregation Module (DFAM) for multi-level feature fusion.
  • Implemented a CoordConv Scheme to enhance feature discrimination.
  • Used Deep Supervision for improved low-level feature accuracy and incorporated Uncertainty Modeling.
  • DFAM significantly outperformed state-of-the-art methods on benchmarks, achieving higher localization accuracy.
  • Improved performance noted in repetitive and sparse texture areas through enhanced feature representation.

Abstract

As scene coordinate regression (SCoRe) methods become prevailing in the area of visual camera localization, the issue of repetitive or sparse texture scenes continues to be a concern. Specifically, they will suffer from performance degeneration due to ambiguous patterns caused by visual similarity. In this work, we propose a novel network for camera localization through a single RGB image, with our key insight that taking only high-level feature maps as input can be difficult for the network to accurately model the regression problem due to ambiguous patterns and utilizing the rich spatial details in low-level feature maps can tackle this issue. The core components of the network are 1) Deep Feature Aggregation Module (DFAM), which eliminates the difference among the different level feature representations and fuses multi-level context information; 2) CoordConv Scheme, which further improves the discrimination of features in repetitive or sparse texture areas of the image; 3) Deep Supervision, which endows low-level feature maps with direct supervision from the ground truth to improve the accuracy of camera localization; 4) Uncertainty Modeling, which quantifies the prediction errors stemming from the intrinsic noise in the data. Moreover, to maximize the power of DFAM, we embed channel attention modules into it to prune redundant and noisy features, through which we can refine the different level feature maps. Our network is designed to be lightweight and efficient, and the proposed DFAM can be integrated into general SCoRe-based networks. Comprehensive experiments demonstrate the effectiveness of DFAM and the superiority of our network over the state-of-the-art methods on two benchmarks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2022) studied this question.

synapsesocial.com/papers/6a1f7e0fccd4fd538e0723cchttps://doi.org/10.1109/lra.2022.3146946
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization2015 · 2,400 citations
  2. 2Hierarchical Scene Coordinate Classification and Regression for Visual\n Localization2019 · 139 citations
  3. 3CamNet: Coarse-to-Fine Retrieval for Camera Re-Localization2019 · 141 citations
  4. 4Regression Forest Based RGB-D Visual Relocalization Using Coarse-to-Fine Strategy2020 · 14 citations
  5. 5KinectFusion: Real-time dense surface mapping and tracking2011 · 4,005 citations