Analysis reports improved accuracy and latency in autonomous driving with deep learning methods, indicating better sensor integration for intelligent transportation systems.
Autonomous vehicles require accurate perception for safe and effective navigation. Relying solely on RGB cameras or LiDAR is unreliable under poor lighting, occlusion, or adverse weather. Traditional multi-sensor fusion methods (early or late fusion) often ignore contextual cues, lack robustness to sensor failure, and suffer from high latency and lower accuracy in complex traffic. To address these issues, this paper proposes a deep learning-based middle fusion framework using EfficientNetV2-S and VoxelNet for feature extraction, and a Transformer encoder for adaptive, attention-based fusion. The model uses environment-aware, modality-specific encoders and dynamically aligns sensor features based on context. It achieved a mean Average Precision of 91.6%, bounding box accuracy of 95.5%, and an inference speed of 41.3 ms/frame. Performance remained strong under night (82.7%) and occlusion (92.1%) conditions, with a Fusion Robustness Index of 0.81. The approach outperformed FCN8, U-Net, and early fusion models in accuracy, speed, and fault tolerance, offering a real-time, robust perception solution for intelligent transportation systems.
No takes yet. Share an insight, caveat, or question.
Wei-min et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: