Typical visual environments contain a rich array of colors, textures, surfaces, and objects, but it is well established that observers do not have access to all of these visual details, even over short intervals (R. A. Rensink, J. K. O'Regan, & J. J. Clark, 1997). Rather, it seems that human vision extracts only partial information from every glance. What is the nature of this selective encoding of the scene? Although there is considerable research on short-term coding of individual objects, much less is known about the representation of a natural scene in visual short-term memory (VSTM). Here, we examine the VSTM of natural scenes using a local recognition task. A major finding is that local recognition performance is better when image segments are viewed in the context of coherent rather than scrambled scenes, suggesting that observers rely on an encoding of a global 'gist' of the scene. Variations on this experiment allow quantification of the role of multiple factors in local recognition. Color statistics and the global configural context are found to be more important than local features of the target, even for a local recognition task.
No takes yet. Share an insight, caveat, or question.
Velisavljevic et al. (2008) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: