PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 11, 20240 citationsOpen Access

Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation

View Full Paper
TSTong ShaoDolby (United States)ZTZhuotao TianHarbin Institute of TechnologyHZHang ZhaoChinese Academy of Tropical Agricultural Sciences

Key Points

Key points are not available for this paper at this time.

Abstract

CLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot capabilities. Despite its success, its application to OVSS faces challenges due to its initial image-level alignment training, which affects its performance in tasks requiring detailed local context. Our study delves into the impact of CLIP's CLS token on patch feature correlations, revealing a dominance of "global" patches that hinders local feature discrimination. To overcome this, we propose CLIPtrase, a novel training-free semantic segmentation strategy that enhances local feature awareness through recalibrated self-correlation among patches. This approach demonstrates notable improvements in segmentation accuracy and the ability to maintain semantic coherence across objects.Experiments show that we are 22.3% ahead of CLIP on average on 9 segmentation benchmarks, outperforming existing state-of-the-art training-free methods.The code are made publicly available at: https://github.com/leaves162/CLIPtrase.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shao et al. (2024) studied this question.

synapsesocial.com/papers/68e60ad1b6db64358759e303https://doi.org/10.48550/arxiv.2407.08268
Ask AI
Helpful
Bookmark
Share
View Full Paper