We investigated mental representations of real road scenes using a prediction task. On each trial, 48 licensed drivers watched a road scene for 2 seconds and immediately after selected what the scene would look like 2 seconds in the future in a five-alternative forced choice (5AFC) task. We factorially manipulated three visual properties of the preview to examine their effects on performance: motion, scene context, and stereoscopic depth. We developed a binocular dashcam video dataset of real road scenes simulating human interpupillary distance. We manipulated motion by displaying still images or video previews. Scene context was manipulated by showing videos recorded on urban roads, which are visually denser, or in highway roads, which are visually sparser. To investigate whether small disparity signals are used in this task, we manipulated stereoscopic depth by varying stimulus presentation: Each eye received video from the corresponding camera, only one eye received video, or both eyes received identical videos. Participants were able to correctly select the future appearance of the road to some extent, with better precision for video than stills and for urban scenes than highway scenes. Stereoscopic cues did not affect performance, indicating that participants likely relied more heavily on monocular depth cues in this task. These results suggest that mental representations of complex natural scenes are temporally imprecise and are consistent with, but do not uniquely require, projections of future states.
Song et al. (Thu,) studied this question.