PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 30, 20241 citations

BEVMamba: Time Sequence Dense Bird's-Eye-View Perception Modeling with State Space Model

View Full Paper
XLXiao LiuChina Aerodynamics Research and Development CenterJZJiaru ZhongBeijing Institute of TechnologyCSChao SunBeijing Institute of Technology

Key Points

Key points are not available for this paper at this time.

Abstract

BEV-based 3D perception with multi-frame images input is crucial for autonomous driving. However, current methods for temporal BEV perception fail to fully utilize long sequence features because of local fusion or high complexity. Recently, Mamba, a powerful temporal modeling network with linear complexity, has shown exceptional performance in various 2D vision tasks, but its application to 3D perception tasks remains unexplored. Therefore, this paper proposes a general BEV perception backbone named BEVMamba, which is the first work to leverage State Space Model for 3D perception. Built upon the BEVFormer, to adapt Mamba for 3D perception we first add Hybrid Positional Encoding to the BEV features, enabling the networks to be aware of their spatial-temporal position. In the Temporal SSM block, the proposed 3D Factorized Scan ensures that historical BEV features are enriched with global temporalspatial information. Subsequently, the Spatial-Temporal Corridor Fusion aggregates all BEV features in a physically meaningful manner, achieving precise feature fusion. The reliable BEV features obtained by BEVMamba are used for various perception tasks, including 3D object detection and 3D occupancy prediction. Results on the nuScenes and Occ-3D nuScenes datasets show that BEVMamba outperforms its baseline BEVFormer in both dense and sparse perception tasks and demonstrates competitive performance compared to other methods, highlighting the potential of Mamba in 3D perception tasks. The code will be available at https://github.com/Liuxiaoaaa/bevmamba.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2024) studied this question.

synapsesocial.com/papers/68e5e808b6db64358757cc7chttps://doi.org/10.36227/techrxiv.172236567.72239502/v1
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1M-BEV: Masked BEV Perception for Robust Autonomous Driving2024 · 12 citations
  2. 2ME$^3$-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception2025
  3. 3Mamba-BEV: A Multiscale State-Space Framework for 3D Object Detection from Point Clouds2026
  4. 4MTC-BEV: Semantic-Guided Temporal and Cross-Modal BEV Feature Fusion for 3D Object Detection2025 · 1 citations
  5. 5Exploring Recurrent Long-Term Temporal Fusion for Multi-View 3D Perception2024 · 74 citations