Local feature matching plays a critical role in robotic SLAM and visual localization. However, in weakly textured indoor industrial environments, lightweight appearance-based methods often struggle to learn discriminative and stable local features. To address this challenge, this paper proposes GAEFeat, short for Geometry-Aware Efficient Feature, a lightweight vision–geometric feature learning network. To address the scarcity of specialized training data, we integrated robotic arm pose priors with depth information to automatically generate cross-view supervision signals and surface-normal labels. Based on this strategy, we constructed two complementary datasets, including a simulated dataset and a real-world dataset, to support feature learning and evaluation in weakly textured indoor industrial environments. For feature extraction, we design a dual enhancement mechanism consisting of a geometric auxiliary branch and a geometry-aware enhancement (GAE) module. The former guides the network to perceive local surface structures through surface normal supervision, while the latter utilizes a gating mechanism to achieve deep fusion between geometric priors and 2D texture descriptors. Experimental results demonstrate that GAEFeat achieves strong robustness and high inference efficiency in relative pose estimation, homography estimation, and visual localization tasks, with particularly notable advantages in near-field, weakly textured industrial scenes. The framework achieves an inference latency of only 3.9 ms on the NVIDIA Jetson AGX Orin edge platform, demonstrating its real-time capability and practical potential for deployment in edge computing environments.
Sun et al. (Sun,) studied this question.