Los puntos clave no están disponibles para este artículo en este momento.
Traditional visual SLAM systems often suffer from localization drift in dynamic environments due to interference from moving objects. Although semantic segmentation and depth-based masking methods have improved performance, they may still suffer from boundary under-segmentation and missed detections due to truncation of dynamic objects. To address these challenges, we propose a cascaded framework, DMSG-SLAM, a cascaded visual SLAM system that fuses Depth-Mask, Semantic information and Geometry constraints for dynamic environments. A lightweight object detection network, combined with depth consistency, is first employed to generate instance-like masks for preliminary dynamic feature removal. Then, a rotation-aware local epipolar geometric filtering mechanism is introduced to suppress residual features near object boundaries and mitigate perceptual blind spots caused by occlusion or truncation. Within potential dynamic regions, the epipolar threshold is adaptively switched according to the estimated inter-frame rotation to provide a more conservative filtering effect under challenging motion conditions. In addition, a TSDF-based dense volumetric map is incorporated to reconstruct more consistent surfaces. Experiments on highly dynamic sequences from the TUM RGB-D dataset indicate that DMSG-SLAM achieves competitive accuracy in dynamic environments, with localization performance improving by up to 90% compared to ORB-SLAM2.
Li et al. (Sun,) studied this question.