PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 11, 20240 citationsOpen Access

Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation

View Full Paper
TSTong ShaoZTZhuotao TianHZHang Zhao

Key Points

Key points are not available for this paper at this time.

Abstract

CLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot capabilities. Despite its success, its application to OVSS faces challenges due to its initial image-level alignment training, which affects its performance in tasks requiring detailed local context. Our study delves into the impact of CLIP's CLS token on patch feature correlations, revealing a dominance of "global" patches that hinders local feature discrimination. To overcome this, we propose CLIPtrase, a novel training-free semantic segmentation strategy that enhances local feature awareness through recalibrated self-correlation among patches. This approach demonstrates notable improvements in segmentation accuracy and the ability to maintain semantic coherence across objects.Experiments show that we are 22.3% ahead of CLIP on average on 9 segmentation benchmarks, outperforming existing state-of-the-art training-free methods.The code are made publicly available at: https://github.com/leaves162/CLIPtrase.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shao et al. (2024) studied this question.

synapsesocial.com/papers/68e60ad1b6db64358759e303https://doi.org/10.48550/arxiv.2407.08268
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation2024
  2. 2CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation2024
  3. 3Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation2025
  4. 4Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation2024
  5. 5TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training2024 · 27 citations