PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 17, 2026Results in Engineering0 citationsOpen Access

ViFusion: Enhanced 3D Object Detection via Virtual Point Augmentation and Dynamic Multi-Source Feature Fusion

View Full Paper
SSShaojing SongZXZhicheng XiangShanghai Polytechnic UniversityMIMuhammad IlyasKazakh-British Technical University

Key Points

  • The aim is to enhance 3D object detection performance through efficient virtual point generation and feature fusion methods.
  • Developed ViFusion framework combining VPA and DMF.
  • VPA generates high-quality virtual points for object detection.
  • DMF facilitates dynamic multi-source feature fusion with optimized resource use.
  • Combined VPA and Ground Truth Augmentation to refine data quality.
  • Tested performance on the KITTI benchmark.
  • Achieved up to 5.84% improvement in Average Precision (AP) on Pointpillars for car detection.
  • Reached 76.65% AP for bicycles using Casa-V integrated with ViFusion.
  • Demonstrated efficient performance increase with reduced computational resources.

Abstract

• ViFusion boosts 3D detection with VPA and DMF on the KITTI benchmark. • VPA generates high-quality virtual points to improve long-range object detection. • DMF enables efficient multi-source feature fusion with low computational cost. • Combined VPA+GT-Aug improves AP by up to 5.84% on Pointpillars (car, moderate). • Achieves 76.65% AP for bicycles on KITTI with Casa-V + ViFusion integration. Due to the sparsity of point clouds, the performance of 3D object detection in autonomous driving is significantly limited. Recent methods employ depth completion networks to generate virtual points for enhancing object representations. However, the high density and noise of these virtual points result in higher resource consumption during training, further exacerbating computational complexity, particularly in multi-source feature fusion scenarios. To address these challenges, an efficient optimization framework, ViFusion (Virtual-enhanced Feature Fusion), is proposed. This framework comprises two key components: VPA (Virtual Point Augmentor) and DMF (Dynamic Multi-Source Feature Fusion). Before feature extraction, VPA enhances key points in dense virtual point clouds through an improved clustering algorithm, efficiently generating high-precision points that significantly enrich object representations. Following feature extraction, DMF dynamically assigns weights to features from multiple sources via a set of pre-designed Multi-Layer Perceptrons (MLP), avoiding complex network designs, thereby enabling efficient fusion while mitigating computational overhead. Compared with existing methods, both modules have shown good results on multiple basic detectors. VPA has excellent accuracy performance and can be effectively combined with the traditional data enhancement method Ground Truth Augmentation (GT-Aug) to achieve better performance. DMF provides excellent performance increase with less resource consumption. The two can be effectively combined, integrating both VPA and DMF into CasA-V (Cascaded attention applied to Voxel-RCNN) which yields the best performance, achieving an AP of 76.65% in detecting medium-difficulty cyclist categories in the KITTI validation set.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Song et al. (2026) studied this question.

synapsesocial.com/papers/69b8ef6ddeb47d591b8c57f1https://doi.org/10.1016/j.rineng.2026.109983
Ask AI
Helpful
Bookmark
Share
View Full Paper