PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 20242 citationsOpen Access

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

View Full Paper
AKAndreas KoukounasGMGeorgios MastrapasMGMichael GüntherGoethe University Frankfurt

Key Points

Key points are not available for this paper at this time.

Abstract

Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors. These models are key to multimodal information retrieval and related tasks. However, CLIP models generally underperform in text-only tasks compared to specialized text models. This creates inefficiencies for information retrieval systems that keep separate embeddings and models for text-only and multimodal tasks. We propose a novel, multi-task contrastive training method to address this issue, which we use to train the jina-clip-v1 model to achieve the state-of-the-art performance on both text-image and text-text retrieval tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Koukounas et al. (2024) studied this question.

synapsesocial.com/papers/68e67aa1b6db643587604f7fhttps://doi.org/10.48550/arxiv.2405.20204
Ask AI
Helpful
Bookmark
Share
View Full Paper