PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 20240 citationsOpen Access

CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models

View Full Paper
SCSantiago CastroAZAmir ZiaiASAvneesh Saluja

Key Points

Key points are not available for this paper at this time.

Abstract

Recent years have witnessed a significant increase in the performance of Vision and Language tasks. Foundational Vision-Language Models (VLMs), such as CLIP, have been leveraged in multiple settings and demonstrated remarkable performance across several tasks. Such models excel at object-centric recognition yet learn text representations that seem invariant to word order, failing to compose known concepts in novel ways. However, no evidence exists that any VLM, including large-scale single-stream models such as GPT-4V, identifies compositions successfully. In this paper, we introduce a framework to significantly improve the ability of existing models to encode compositional language, with over 10% absolute improvement on compositionality benchmarks, while maintaining or improving the performance on standard object-recognition and retrieval benchmarks. Our code and pre-trained models are publicly available at https://github.com/netflix/clove.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Castro et al. (2024) studied this question.

synapsesocial.com/papers/68e781fab6db6435876f5735https://doi.org/10.48550/arxiv.2402.15021
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition2024
  2. 2Iterated Learning Improves Compositionality in Large Vision-Language Models2024
  3. 3Semantic Compositions Enhance Vision-Language Contrastive Learning2024
  4. 4Do Vision-Language Models Understand Compound Nouns?2024
  5. 5Evaluating Compositional Generalisation in VLMs and Diffusion Models2025