Key points are not available for this paper at this time.
Although Vision-Language Models (VLMs) have demonstrated great potential, limited integration of spatial context via textual input has constrained their performance in geospatial analysis, particularly for location-based services (LBS). To this end, we propose Spatial-context Prompt Tuning, tailored to global image geo-localization tasks. Six spatial-context dimensions (i.e., geospatial image types, geo-localization clues, spatial patterns, land use/land cover, urban perception, and urban development) are designed to create Visual Question Answering (VQA) prompts with GPT-4, based on three imagery types (street view images, satellite images, map tiles) sampled from 790 populous cities. Next, we evaluate the efficacy of these dimensions based on a leading open-source VLM – Contrastive Language-Image Pre-training (CLIP). Results demonstrate consistent, task-dependent accuracy improvements in image geo-localization: land-use prompts improve city-level accuracy by 4.5%, and image-type prompts increase country-level accuracy by 6.7% and continent-level accuracy by 3.0%. Our key contributions include: (1) revealing how spatial versus non-spatial context affects prompt-tuned VLM performance, (2) designing reusabGle six dimensions of spatial context that support explainable, context-aware VLMs for geospatial applications, and (3) enhancing geo-localization accuracy across heterogeneous geospatial imagery. This work charts a positive direction for Geospatial Artificial Intelligence (GeoAI)-empowered research, enabling more effective and interpretable VLM applications in geospatial domains.
Wu et al. (Sun,) studied this question.