Methods for estimating building height based on street view images encounter challenges such as dependence on labeled data, image distortion, terrain relief, and occlusion. To address these issues, this study proposes a zero-shot segmentation model and a multi-feature constraint framework based on segment anything model 2 (SAM2) for arbitrary object segmentation without training. Two building semantic constraints are introduced: geospatial constraints, wherein SAM2 building prompt points are generated by intersecting the street-view skyline with building footprint orientation; and visual feature constraints, in which vegetation index and texture features are fused to eliminate non-building feature points, suppressing occlusion interference. Street view shooting location projections onto building footprints replace roof corner points, reducing equirectangular projection distortion. During height inversion, digital elevation model data convert relative heights into absolute elevations, whereas multi-view consistency optimization resolves errors from single-view occlusion and terrain relief. For buildings with insufficient street-view coverage, elevation interpolation leverages neighborhood spatial correlation. Final height is derived from roof–ground elevation difference. Experiments conducted in Nanjing (52 buildings) yielded an average absolute error of 1.36 m, with 100% accuracy within 4 m. A large-scale experiment in Guangzhou (>20,000 buildings) further demonstrated the superior accuracy, robustness, and scalability of the proposed framework.
Lu et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: