PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 16, 202426 citationsOpen Access

Transductive Zero-Shot and Few-Shot CLIP

View Full Paper
SMSégolène MartinUniversité Libre de BruxellesYHYunshi HuangBeijing Academy of Artificial IntelligenceFSFereshteh ShakeriLiv Hospital

Key Points

Key points are not available for this paper at this time.

Abstract

Transductive inference has been widely investigated in few-shot image classification, but completely overlooked in the recent, fast growing literature on adapting vision-langage models like CLIP. This paper addresses the transductive zero-shot and few-shot CLIP classification challenge, in which inference is performed jointly across a mini-batch of unlabeled query samples, rather than treating each instance independently. We initially construct informative vision-text probability features, leading to a classification problem on the unit simplex set. Inspired by Expectation-Maximization (EM), our optimization-based classification objective models the data probability distribution for each class using a Dirichlet law. The minimization problem is then tackled with a novel block Majorization-Minimization algorithm, which simultaneously estimates the distribution parameters and class assignments. Extensive numerical experiments on 11 datasets underscore the benefits and efficacy of our batch inference approach.On zero-shot tasks with test batches of 75 samples, our approach yields near 20% improvement in ImageNet accuracy over CLIP's zero-shot performance. Additionally, we outperform state-of-the-art methods in the few-shot setting. The code is available at: https://github.com/SegoleneMartin/transductive-CLIP.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Martin et al. (2024) studied this question.

synapsesocial.com/papers/68e64877b6db6435875d96c7https://doi.org/10.1109/cvpr52733.2024.02722
Ask AI
Helpful
Bookmark
Share
View Full Paper