PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 11, 2026Scientific Reports0 citationsOpen Access

PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss

SBSheethal BhatSiemens (Germany)AMAwais MansoorSiemens (United States)BGBogdan GeorgescuSiemens (United States)

Key Points

  • The aim is to improve fine-grained spatial understanding in vision-language models for medical image analysis.
  • Introduced Patch-CLIP framework using contrastive loss for aligning image patch-level and text embeddings.
  • Utilized two Chest X-ray datasets for evaluating model performance across multiple abnormality detection tasks.
  • Achieved state-of-the-art performance by focusing on local patch-level features rather than just global representations.
  • Patch-CLIP outperformed traditional models across eight abnormality detection tasks with state-of-the-art metrics.
  • Reduced false positives while maintaining comparable sensitivity to standard saliency-based methods.
  • Provided enhanced localization and interpretability of key findings in medical images.

Abstract

Abstract Vision-Language (VL) models such as Contrastive Language-Image pretraining (CLIP) have shown remarkable zero-shot classification capabilities by jointly learning from large-scale image–text datasets using multimodal self-supervised learning (SSL). However, while these models capture strong global semantics, they often struggle with fine-grained spatial understanding, thereby limiting their effectiveness in downstream tasks like object detection and medical abnormality localization 2 . To address this limitation, we propose Patch-CLIP, a novel VL framework that introduces a contrastive loss aligning image patch-level embeddings with text embeddings. Unlike conventional VL approaches that only leverage global image representations, our method utilizes local patch-level features to encode spatial context, enabling effective learning of localization cues. Applied to two Chest X-ray (CXR) datasets, Patch-CLIP achieves state-of-the-art (SOTA) performance across eight abnormality detection tasks. Furthermore, the resulting patch prediction maps substantially reduce false positives at comparable sensitivity levels compared to standard saliency-based methods, providing more precise and interpretable localization of key findings. The code is available at https://github.com/Siemens-Healthineers/patch-clip

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bhat et al. (2026) studied this question.

synapsesocial.com/papers/6a0172233a9f334c2827243ehttps://doi.org/10.1038/s41598-026-52235-x
Ask AI
Helpful
Bookmark
Share
View Full Paper