PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 19, 2025Journal of Imaging2 citationsOpen Access

Semantic-Enhanced and Temporally Refined Bidirectional BEV Fusion for LiDAR–Camera 3D Object Detection

View Full Paper
XQXinying QuQKQin KaiYLYa-Ping Li

Key Points

  • SETR-Fusion improved detection accuracy, achieving 71.2% mAP and 73.3% NDS on the nuScenes test set.
  • Key components include the Discriminative Semantic Saliency Activation and Bilateral Cross-Attention Fusion modules.
  • By enhancing multimodal information interaction, the method effectively addresses limitations in LiDAR detection.
  • The approach leverages semantic features from images and point clouds for robust environmental perception.

Abstract

In domains such as autonomous driving, 3D object detection is a key technology for environmental perception. By integrating multimodal information from sensors such as LiDAR and cameras, the detection accuracy can be significantly improved. However, the current multimodal fusion perception framework still suffers from two problems: first, due to the inherent physical limitations of LiDAR detection, the number of point clouds of distant objects is sparse, resulting in small target objects being easily overwhelmed by the background; second, the cross-modal information interaction is insufficient, and the complementarity and correlation between the LiDAR point cloud and the camera image are not fully exploited and utilized. Therefore, we propose a new multimodal detection strategy, Semantic-Enhanced and Temporally Refined Bidirectional BEV Fusion (SETR-Fusion). This method integrates three key components: the Discriminative Semantic Saliency Activation (DSSA) module, the Temporally Consistent Semantic Point Fusion (TCSP) module, and the Bilateral Cross-Attention Fusion (BCAF) module. The DSSA module fully utilizes image semantic features to capture more discriminative foreground and background cues; the TCSP module generates semantic LiDAR points and, after noise filtering, produces a more accurate semantic LiDAR point cloud; and the BCAF module’s cross-attention to camera and LiDAR BEV features in both directions enables strong interaction between the two types of modal information. SETR-Fusion achieves 71.2% mAP and 73.3% NDS values on the nuScenes test set, outperforming several state-of-the-art methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qu et al. (2025) studied this question.

synapsesocial.com/papers/68d466a831b076d99fa64e7dhttps://doi.org/10.3390/jimaging11090319
Ask AI
Helpful
Bookmark
Share
View Full Paper