Vision foundation models (VFMs) enable high-quality 2D instance masks, yet their lifted pseudo-point clouds suffer from scale ambiguity, structural noise, and temporal inconsistency, limiting their utility in 3D annotation. Existing automatic labeling methods either rely on expensive light detection and ranging (LiDAR) sensors or fail to enforce physical plausibility in dynamic roadside scenes. This study proposes a LiDAR-free radar–visual auto-labeling framework that leverages cross-modal spatio-temporal consistency between millimeter-wave radar trajectories and visual pseudo-point clouds to self-correct 3D geometry. The method first associates radar points, 2D masks, and pseudo-point clouds into object-centric sequences. Then, an uncertainty-aware pose fusion module combines motion-derived and structure-derived orientations using automatically solved road priors. Finally, the pseudo-point cloud is refined in canonical space by optimizing stable semantic landmarks from temporally consistent masks and propagating their corrections globally. Evaluated on a real-world roadside dataset, the method achieves 49.1% bird’s-eye-view (BEV) intersection over union (IoU) and 43.0% 3D IoU, outperforming a radar–camera fusion baseline by 5.5/5.9 points. Downstream experiments further show that the generated pseudo-labels and semantic enhancement are useful under the evaluated detector configurations, while broader validation remains future work.
Zhu et al. (Fri,) studied this question.