Literature review demonstrates expanding accuracy and utility of monocular depth estimation for volunteered street view imagery, indicating high potential for low-cost 3D urban environment analysis.
The utilisation of computer vision in urban studies has become common practice due to its capacity to diminish the financial burden associated with field surveys. Monocular Depth Estimation (MDE) is a recent branch of computer vision that has been shown to be capable of predicting three-dimensional information from a single image. Street View Imagery (SVI) refers to the collection of a substantial dataset comprising urban images. The purpose of this paper is to analyse the potential of applying MDE to Volunteered SVI (VSVI) in urban studies. Following the processes of acquisition and screening, a total of 102 MDE and 42 studies employing SVI are utilised to delineate this potential association. The number of MDE models, training strategies and the volume of training, validation and evaluation datasets have all increased over the years. Notably, MDE models have significantly improved in accuracy through the KITTI benchmark test. Despite the gap between MDE developers and urban study practitioners, as well as the misuse of VSVI, the application of MDE in SVI-based urban studies is substantial. Based on observable potentials, this study further proposes a future framework for using MDE in VSVI-based urban studies, which can contribute to the field of computer vision in a built environment.
No takes yet. Share an insight, caveat, or question.
Nguyen et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: