PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 20260 citationsOpen Access

VLMs using language-guided inference capture context-sensitivity of human object recognition behavior

View Full Paper
KRKarim RajaeiRCRadoslaw Martin CichyHSHamid Soltanian-Zadeh

Key Points

  • The research aims to investigate the impact of structured scene context on object recognition and how well artificial vision models replicate this human behavior.
  • Conducted human behavioral experiments with 3D simulated indoor scenes.
  • Manipulated contextual coherence by using intact and phase-scrambled scenes.
  • Compared performance of conventional visual models (CNNs, ViTs) with vision-language models (VLMs).
  • Humans showed better object recognition in coherent scenes, especially under difficult conditions.
  • Traditional visual models did not replicate human context sensitivity.
  • VLMs, particularly those trained with language, approached human-like accuracy in recognition.

Abstract

Human vision is a context-sensitive process that interprets objects in relation to their surroundings. While behavioral research has long shown that scene context facilitates object recognition, the underlying computational mechanisms—and the extent to which artificial vision models replicate this ability—remain unclear. Here, we addressed this gap by combining human behavioral experiments with computational modeling to investigate how structured scene context influences object recognition. Using a 3D simulation framework, we embedded target objects into indoor scenes, and manipulated contextual coherence between objects and scenes by using either intact scenes or their phase-scrambled versions as context. Humans showed a robust object recognition advantage in coherent scenes, particularly under challenging conditions such as occlusion, crowding, or non-canonical viewpoints. Conventional vision models—including convolutional neural networks (CNNs) and vision transformers (ViTs)—failed to replicate this effect. In contrast, vision-language models (VLMs), particularly those using ViT architectures and trained with language supervision (e.g., CLIP), approached human-like accuracy. This shows that semantically rich and category-structured representations are required for modelling context sensitivity. Notably, context sensitive behavior was closest to humans in VLMs when using language-guided inference at test time. This suggests that how a model accesses its representations during inference is relevant for enabling context-sensitive behavior. Together, this work offers steps towards a computational account of contextual facilitation of objects by scenes, and highlights zero-shot inference as an interesting alignment metric when benchmarking artificial and biological vision.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rajaei et al. (2026) studied this question.

synapsesocial.com/papers/69e3213840886becb6540692https://doi.org/10.17169/refubium-51941
Ask AI
Helpful
Bookmark
Share
View Full Paper