Introduction We previously investigated color constancy in photorealistic virtual reality (VR) and developed a Deep Neural Network (DNN) that predicts reflectance from rendered images. Methods We combine both approaches to compare and study a model and human performance with respect to established color constancy mechanisms: local surround, maximum flux and spatial mean. Rather than evaluating the model against physical ground truth, model performance was assessed using the same achromatic object selection task employed in the human experiments. The model, a ResNet based U-Net from our previous work, was pre-trained on rendered images to predict surface reflectance. We then applied transfer learning, fine-tuning only the network's decoder on images from the baseline VR condition. To parallel the human experiment, the model's output was used to perform the same achromatic object selection task across all conditions. Results A strong correspondence between the model and human behavior was observed. Both achieved high constancy under baseline conditions and showed similar, condition-dependent performance declines when the local surround or spatial mean color cues were removed. Discussion These results show that a pixel-wise DNN trained on naturalistic image statistics can reproduce the structure of human color constancy behavior across controlled cue manipulations, supporting the view that human constancy can arises from the integration of multiple scene-based cues without explicit illuminant estimation.
Heidari-Gorji et al. (Tue,) studied this question.