PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 18, 2026IEEE Transactions on Image Processing1 citations

DiCLIP: Diffusion Model Enhances CLIP’s Dense Knowledge for Weakly Supervised Semantic Segmentation

View Full Paper
ZYZhiwei YangPSPengfei SongYMYucong Meng

Key Points

  • The aim is to improve weakly supervised semantic segmentation by enhancing CLIP's capabilities using a diffusion model.
  • Proposed DiCLIP framework utilizes Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules.
  • Implemented an Attention Clustering Refinement (ACR) module to optimize correlation map extraction from the diffusion model.
  • Applied a dynamic key-value cache model to enrich text embeddings for better semantic representation.
  • DiCLIP outperforms state-of-the-art methods on PASCAL VOC and MS COCO datasets.
  • Significant reduction in training costs compared to previous approaches.
  • Enhanced pixel-level predictions demonstrated through improved CAM generation.

Abstract

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced to generate CAMs in WSSS. However, previous WSSS methods solely adopt CLIP's vision-language paired property for dense localization, neglecting its inherently limited dense knowledge across both visual and text modalities, which renders CAM generation suboptimal. In this work, we propose DiCLIP, a novel WSSS framework that leverages the generative diffusion model to enhance CLIP's dense knowledge across two modalities. Specifically, Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules are proposed for dense prediction enhancement. To improve the spatial awareness of visual features, our VCE module utilizes diffusion's reliable spatial consistency to mitigate the over-smoothing issue in CLIP's attention. It designs the Attention Clustering Refinement (ACR) module to reliably extract diverse correlation maps from the diffusion model. The correlation maps act as a diversity bias for CLIP's self-attention, recursively pushing its visual features towards a more discriminative dense distribution. To augment the semantics of text embeddings, our TSA module argues that a single text modality is insufficient to encompass the variability of visual categories. Thus, we leverage diffusion's generative power to maintain a dynamic key-value cache model, shifting CAM generation from a patch-text matching mechanism to a novel visual knowledge retrieval paradigm.With these enhancements, DiCLIP not only outperforms state-of-the-art methods on PASCAL VOC and MS COCO but also significantly reduces training costs. Code will be publicly available at https://github.com/zwyang6/DiCLIP.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2026) studied this question.

synapsesocial.com/papers/6a0aac6d5ba8ef6d83b6fc7ehttps://doi.org/10.1109/tip.2026.3692055
Ask AI
Helpful
Bookmark
Share
View Full Paper