PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 11, 2026Scientific Reports0 citationsOpen Access

PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss

SBSheethal BhatAMAwais MansoorBGBogdan Georgescu

Key Points

  • The aim is to improve fine-grained spatial understanding in vision-language models for medical image analysis.
  • Introduced Patch-CLIP framework using contrastive loss for aligning image patch-level and text embeddings.
  • Utilized two Chest X-ray datasets for evaluating model performance across multiple abnormality detection tasks.
  • Achieved state-of-the-art performance by focusing on local patch-level features rather than just global representations.
  • Patch-CLIP outperformed traditional models across eight abnormality detection tasks with state-of-the-art metrics.
  • Reduced false positives while maintaining comparable sensitivity to standard saliency-based methods.
  • Provided enhanced localization and interpretability of key findings in medical images.

Abstract

Abstract Vision-Language (VL) models such as Contrastive Language-Image pretraining (CLIP) have shown remarkable zero-shot classification capabilities by jointly learning from large-scale image–text datasets using multimodal self-supervised learning (SSL). However, while these models capture strong global semantics, they often struggle with fine-grained spatial understanding, thereby limiting their effectiveness in downstream tasks like object detection and medical abnormality localization 2 . To address this limitation, we propose Patch-CLIP, a novel VL framework that introduces a contrastive loss aligning image patch-level embeddings with text embeddings. Unlike conventional VL approaches that only leverage global image representations, our method utilizes local patch-level features to encode spatial context, enabling effective learning of localization cues. Applied to two Chest X-ray (CXR) datasets, Patch-CLIP achieves state-of-the-art (SOTA) performance across eight abnormality detection tasks. Furthermore, the resulting patch prediction maps substantially reduce false positives at comparable sensitivity levels compared to standard saliency-based methods, providing more precise and interpretable localization of key findings. The code is available at https://github.com/Siemens-Healthineers/patch-clip

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bhat et al. (2026) studied this question.

synapsesocial.com/papers/6a0172233a9f334c2827243ehttps://doi.org/10.1038/s41598-026-52235-x
Ask AI
Helpful
Bookmark
Share
View Full Paper