Randomized trial evaluates metric distance prediction for vehicles and pedestrians, suggesting effective alternatives to traditional systems.
Monocular vision-based Advanced Driver Assistance Systems (ADAS) offer a cost-effective and scalable alternative to LiDAR and Radar-based systems, with a simpler setup than binocular systems. However, accurately estimating metric distance from monocular images remains challenging, especially under varying environmental conditions. This work introduces a two-stage framework for precise vehicle and pedestrian distance estimation. Instead of relying on dataset-specific models, the framework employs a zero-shot depth estimation model. In the first stage, the framework leverages the pre-trained Depth Anything V2 model for depth map generation and YOLOv11 (trained on BDD100K) for object detection. These models produce relative depth maps and bounding box details with class labels for vehicles and pedestrians. These outputs are combined to extract object-specific features such as bounding boxes and statistical depth features. In the second stage, a Long Short-Term Memory (LSTM)-based distance estimator is trained to predict the metric distance of detected objects using these extracted features. The framework is initially trained on the KITTI 3D object detection dataset and fine-tuned on Lyft Level 5, nuScenes, and ONCE. It achieves competitive performance on KITTI with a Mean Absolute Error (MAE) of 1.18 and a Root Mean Square Error (RMSE) of 2.03. Fine-tuning improves performance on Lyft Level 5 and nuScenes while maintaining KITTI accuracy without catastrophic forgetting. Mixed-dataset training with equal proportions from all four datasets achieves an MAE of 2.74 and RMSE of 4.85 on the mixed validation set, demonstrating strong generalization across diverse driving environments. The framework is further evaluated on the Waymo Open Dataset under deployment-mimicking conditions using detector-derived bounding boxes, confirming practical deployment readiness.
No takes yet. Share an insight, caveat, or question.
Kumar et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: