Vehicle-mounted smartphones are used to increase the frequency of road inspections. If images taken at different times can be accurately matched, temporal changes in road conditions become more apparent for damage analysis. However, the low quality of smartphone sensors, road deterioration, and lighting variations make matching unreliable. Existing matchers generalize poorly to road scenes and are usually trained with 3D supervision, such as depth maps and camera poses, which are unavailable for smartphones. To address this limitation, a self-supervised adaptation with pixel-level precision is proposed, in which the supervision of an existing dense matching model is redesigned to enable training solely on RGB images. A self-supervised labeling pipeline eliminates manual annotation of point correspondences. GPS ensures coarse alignment and bird’s-eye-view transformation mitigates perspective distortion. Then multiple open-source matchers are ensembled to propose correspondences, which are refined by robust fitting and propagated to unlabeled regions. Over five years, 76,882 image pairs were collected across 20 regions in Japan and China. The proposed method boosts correspondence coverage from 30.5% to 82.2% with 98.6% accuracy, maintaining robustness across smartphones, conventional digital cameras, and line-scan cameras.
Xue et al. (2026) studied this question.