PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025ISPRS annals of the photogrammetry, remote sensing and spatial information sciences2 citationsOpen Access

Advancing Mixed Land Use Detection by Embedding Spatial Intelligence into Vision-Language Models

View Full Paper
MWMeiliu WuQHQunying HuangSGSong Gao

Key Points

  • GeospatialCLIP shows superior accuracy in detecting urban mixed land use compared to traditional models and state-of-the-art methods.
  • Utilizing techniques like spatial-context aware prompt engineering, the framework delivers robust multimodal representations.
  • Spatial prompts that provide city-specific cues can significantly enhance detection accuracy.
  • The study underscores the importance of spatial intelligence in advancing the performance of vision-language models.

Abstract

Abstract. Embedding spatial intelligence into vision-language models (VLMs) has offered a promising avenue to improve geospatial decision-making in complex urban environments. In this work, we propose a novel framework that augments the architecture of Contrastive Language-Image Pretraining (CLIP) with the techniques of spatial-context aware prompt engineering and spatially explicit contrastive learning. By leveraging a diverse set of geospatial imagery (e.g., street view, satellite, and map tile images), paired with contextual geospatial text generated and curated via GPT-4, our approach constructs robust multimodal representations that capture visual, textual, and spatial insights. The proposed model, termed GeospatialCLIP, is specifically evaluated for urban mixed land use detection, a critical task for sustainable urban planning and smart city development. Results demonstrate that GeospatialCLIP consistently outperforms traditional vision-based few-shot models (e.g., ResNet-152, Vision Transformers) and exhibits competitive performance with state-of-the-art models such as GPT-4. Notably, the incorporation of spatial prompts, especially those providing city-specific cues, significantly boosts detection accuracy. Our findings highlight the pivotal role of spatial intelligence in refining VLM performance and provide novel insights into the integration of geospatial reasoning within multimodal learning. Overall, this work establishes a foundation for future spatially explicit AI development and applications, paving the way for more comprehensive and interpretable models in urban analytics and beyond.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68d469c831b076d99fa6673dhttps://doi.org/10.5194/isprs-annals-x-4-w7-2025-121-2025
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1GeoPriorclip: a foundational remote sensing vision-language model enhanced with cascaded geographic information priors2026
  2. 2ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization2024 · 1 citations
  3. 3Zero-shot urban function inference with street view images through prompting a pretrained vision-language model2024 · 45 citations
  4. 4Spatial intelligence in vision-language models: a comprehensive survey2026
  5. 5Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning2025