Recent advances in remote sensing have increasingly emphasized multi-view vision, which integrates complementary viewpoints to deliver more complete scene understanding and effectively alleviate occlusion and limited fields of view in crowded environments. In particular, aerial imagery captured by drones provides holistic scene coverage, whereas ground-level cameras offer precise and fine-grained object details. Despite these advantages, large-scale multi-view datasets that jointly incorporate aerial and ground-level perspectives remain scarce, largely due to the practical difficulties of coordinating paired aerial and ground platforms. To overcome this challenge, we develop a ground–aerial camera system that emulates drone viewpoints and, based on this system, construct a large-scale synthetic dataset for aerial–ground multi-view person association. Leveraging this dataset, we propose a novel graph-constrained framework that enforces robust and globally consistent associations across aerial and ground views. Additionally, we introduce an aerial-view-guided people-number estimation module to provide a scene-level constraint for identity association. Extensive experimental results demonstrate that our method consistently outperforms state-of-the-art baselines in multi-view labeling across varying crowd densities.
Zhang et al. (Thu,) studied this question.