Key points are not available for this paper at this time.
We present GraphCLIP, a novel contrastive learning framework for multimodal artwork classification that integrates visual and contextual information to improve predictive accuracy and interpretability. Traditional computer vision methods often fall short in visual arts, where context is crucial. GraphCLIP leverages image data and a Knowledge Graph to extract features from both perspectives. Evaluated on the A r t G r a p h dataset, with over 100,000 artworks in 32 styles and 18 genres, GraphCLIP outperforms existing models in single-task (up to + 8 % in F1-score) and multi-task settings (up to + 6 % ), demonstrating robustness even with unseen classes. Additionally, visual and contextual qualitative explanations enhance model transparency. The versatility of GraphCLIP extends beyond art classification: its methodology can be adapted to other domains where integrating diverse data types is essential. (The code is publicly available at: https://github.com/CILAB-ArtGraph/graphclip.git .) • We introduce GraphCLIP, a contrastive learning framework for artwork classification. • GraphCLIP combines visual data with contextual knowledge. • We achieve state-of-the-art performance on the A r t G r a p h dataset. • We demonstrate robustness with unseen classes in distribution shift scenarios. • We provide visual and contextual explanations to enhance model interpretability.
Scaringi et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: