To address the challenges of unmanned aerial vehicle (UAV) navigation in weak global navigation satellite system (GNSS) environments, this study proposes a novel multimodal feature fusion framework for real-time positioning using a priori high-resolution satellite imagery. This framework utilizes georeferenced satellite images as matching sources and employs a “Multimodal features + LightGlue” algorithm to achieve high-precision cross-modal matching. By combining point, line, and plane features for enhanced robustness in low-texture scenarios, the system further integrates LightGlue’s lightweight confidence classifier to accelerate inference while maintaining high accuracy on challenging image pairs. Consequently, the proposed method outperforms LoFTR, RoMa, SuperPoint + SuperGlue, and SuperPoint + LightGlue in matching performance. Experimental results demonstrate that at a flight altitude of 80 m, the average real-time positioning error is 0.73 m, which increases to 6.24 m at 480 m. Factors such as ground object type, seasonal changes, flight altitude, and satellite image scale significantly influence accuracy. This research demonstrates that the visual navigation system meets practical operational needs for real-time UAV positioning in GNSS-deprived environments.
He et al. (Mon,) studied this question.