PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 19, 2026Nature Biomedical Engineering0 citationsOpen Access

An explainable biomedical foundation model via large-scale concept-enhanced vision–language pretraining

View Full Paper
YNYuxiang NieSHSunan HeYBYequan Bie

Key Points

  • To develop and evaluate ConceptCLIP, an explainable biomedical foundation model capable of delivering state-of-the-art diagnostic accuracy alongside human-interpretable visual and conceptual explanations.
  • Curated MedConcept-23M, a dataset containing 23 million biomedical image–text–concept triplets across diverse modalities.
  • Pretrained ConceptCLIP using joint image–text and region–concept alignment strategies to ground visual features into interpretable concepts.
  • Evaluated performance on a benchmark of 78 datasets spanning 10 imaging modalities and conducted a multi-modality user study with practicing clinicians.
  • ConceptCLIP achieved superior diagnostic accuracy compared to existing foundation models across 10 distinct medical imaging modalities.
  • In a clinical reader study across three modalities, concept-based explanations enabled clinicians to verify model predictions and detect potential diagnostic errors.

Abstract

Artificial intelligence for medical imaging is required to be accurate and interpretable to clinicians. However, current multimodal biomedical foundation models often prioritize performance over explainability. Here we present ConceptCLIP, an explainable biomedical foundation model that achieves state-of-the-art diagnostic accuracy while delivering human-interpretable explanations across diverse imaging modalities. We curate MedConcept-23M, a large-scale dataset comprising 23 million biomedical image–text–concept triplets. Leveraging this dataset, we pretrain ConceptCLIP via joint image–text and region–concept alignment for precise and interpretable medical image analysis. Across a large-scale benchmark covering 78 datasets in 10 imaging modalities, ConceptCLIP demonstrates superior diagnostic performance while providing human-understandable explanations. In a clinician user study spanning three modalities, the concept-based explanations provided by ConceptCLIP help clinicians verify model predictions and identify potential errors. As an explainable biomedical foundation model, ConceptCLIP represents a critical milestone towards the widespread clinical adoption of AI, thereby advancing trustworthy AI in medicine. ConceptCLIP is an explainable foundation model that leverages MedConcept-23M, a large-scale dataset comprising 23 million biomedical image–text–concept triplets, for precise and interpretable medical image analysis and diagnostics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nie et al. (2026) studied this question.

synapsesocial.com/papers/6a85638803308d306e2d6b6bhttps://doi.org/10.1038/s41551-026-01764-x
Ask AI
Helpful
Bookmark
Share
View Full Paper