PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 27, 2026Pattern Recognition3 citationsOpen Access

DMSAA-SLAM: RGB-D SLAM for Dynamic Scenes via Diffusion Self-Attention

View Full Paper
LXLei XiaXLXin LiZWZiyang Wang

Key Points

  • The aim is to improve RGB-D SLAM performance in dynamic scenes using a diffusion model for object detection.
  • Developed a self-attention aggregation module using a diffusion model.
  • Integrated high-precision masks for dynamic tracking in RGB-D SLAM.
  • Validated method accuracy and efficiency on dynamic simulation datasets.
  • Achieved superior accuracy in dynamic scene mapping compared to traditional methods.
  • Effectively reduced tracking errors caused by moving objects.
  • Improved robustness of SLAM process in environments with dynamic features.

Abstract

• Design a self-attention aggregation module using a pre-trained diffusion model. • Integrate high-precision masks into RGB-D SLAM for robust dynamic tracking. • Validate superior accuracy and efficiency on dynamic simulation datasets. In dynamic environments, performing RGB-D SLAM (Simultaneous Localization and Mapping) faces significant challenges primarily due to the presence of moving objects. The motion of these objects can introduce tracking errors and inaccuracies in map construction, thereby compromising the stability and overall performance of the system. To maintain high-precision localization and mapping under such conditions, a SLAM system must effectively detect and handle dynamic objects. To address these challenges, this paper presents a novel RGB-D SLAM method, referred to as DMSAA-SLAM (Dynamic Scene SLAM Based on Diffusion Model Self-Attention Aggregation). The core idea is to leverage a pre-trained stable diffusion model, particularly its self-attention layers, to handle the complexity of dynamic scenes. By employing a multi-resolution aggregation approach, combined with iterative merging and nonmaximum suppression, the proposed method generates high-precision segmentation masks. These masks enable fine-grained segmentation of moving objects and effectively eliminate dynamic feature points, thereby mitigating the impact of dynamic elements on the SLAM process and ensuring efficient and accurate tracking and mapping.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xia et al. (2026) studied this question.

synapsesocial.com/papers/69c6210b15a0a509bde1996bhttps://doi.org/10.1016/j.patcog.2026.113576
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras2017 · 6,195 citations
  2. 2Strong-SLAM: real-time RGB-D visual SLAM in dynamic environments based on StrongSORT2024 · 4 citations
  3. 3FAST-LiDAR-SLAM: A Robust and Real-Time Factor Graph for Urban Scenarios With Unstable GPS Signals2024 · 14 citations
  4. 4Vision-Motion Codesign for Low-Level Trajectory Generation in Visual Servoing Systems2023 · 244 citations
  5. 5DVN-SLAM: Dynamic Visual Neural Slam Based on Local-Global Encoding2025 · 2 citations