Key points are not available for this paper at this time.
Multimodal systems and Large Language Models have shown remarkable capabilities in text-based reasoning, yet their capacity to perceive and interpret visual art remains uncertain. This study examines how CLIP “sees” and understands artworks by comparing their responses to human- and AI-generated paintings in the European tradition from the Renaissance onward. The analysis focuses on its ability to identify style, period and cultural context, as well as potential biases in its perception, evaluated against human judgments.
Asperti et al. (Wed,) studied this question.