The localization performance of visual–inertial simultaneous localization and mapping (VI-SLAM) strongly depends on front-end feature matching. In degraded scenes with low illumination, repetitive textures, and weak textures, traditional geometric front ends often suffer from sparse features and mismatches, resulting in unstable state estimation. To address this issue, this paper proposes Area-to-Point Matching Visual–Inertial SLAM (A2PM-VINS), a visual–inertial SLAM method based on Area-to-Point matching. The method introduces Area-to-Point hierarchical matching and a kinematic temporal inheritance mechanism to improve matching reliability and track continuity, and further designs an Anchor–Explorer feature selection strategy to retain features with higher geometric value for back-end optimization. In addition, a Sub-Window Consistency (SWC) weighting strategy is incorporated into the back end to suppress geometrically deceptive observations with poor temporal continuity and geometric consistency. Experiments on the European Robotics Challenge Micro Aerial Vehicle (EuRoC MAV) dataset show that A2PM-VINS achieves superior or competitive localization accuracy on multiple challenging sequences. The absolute trajectory errors on MH₀4 and MH₀5 are 0. 0983 m and 0. 1191 m, respectively, and stable tracking is maintained on V2₀2, where VINS-Fusion fails. These results show that the proposed method effectively improves the robustness of visual–inertial state estimation in complex degraded environments.
Ma et al. (Wed,) studied this question.