PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 2025npj Digital Medicine24 citationsOpen Access

Large-vocabulary segmentation for medical images with text prompts

View Full Paper
ZZZiheng ZhaoBeijing Institute of TechnologyYZYao ZhangChongqing UniversityCWChaoyi WuShanghai Jiao Tong University

Key Points

  • SAT-Pro achieves a +7.1% average improvement in Dice Similarity Coefficient compared to existing models.
  • The model processes over 22K 3D scans, classifying them into 497 anatomical categories derived from 6502 medical terminologies.
  • Using contrastive learning, the method injects medical knowledge into a text encoder for better segmentation outcomes.
  • SAT-Pro's performance on external datasets shows a +3.7% average improvement over established baselines.

Abstract

This paper aims to build a model that can Segment Anything in 3D medical images, driven by medical terminologies as Text prompts, termed as SAT. Our main contributions are three-fold: (i) We construct the first multimodal knowledge tree on human anatomy, including 6502 anatomical terminologies; Then, we build the largest and most comprehensive segmentation dataset for training, collecting over 22K 3D scans from 72 datasets, across 497 classes, with careful standardization on both image and label space; (ii) We propose to inject medical knowledge into a text encoder via contrastive learning and formulate a large-vocabulary segmentation model that can be prompted by medical terminologies in text form. (iii) We train SAT-Nano (110M parameters) and SAT-Pro (447M parameters). SAT-Pro achieves comparable performance to 72 nnU-Nets—the strongest specialist models trained on each dataset (over 2.2B parameters combined)—over 497 categories. Compared with the interactive approach MedSAM, SAT-Pro consistently outperforms across all 7 human body regions with +7.1% average Dice Similarity Coefficient (DSC) improvement, while showing enhanced scalability and robustness. On 2 external (cross-center) datasets, SAT-Pro achieves higher performance than all baselines (+3.7% average DSC), demonstrating superior generalization ability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2025) studied this question.

synapsesocial.com/papers/68bb4e016d6d5674bcd026c9https://doi.org/10.1038/s41746-025-01964-w
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SALT: Introducing a Framework for Hierarchical Segmentations in Medical Imaging using Softmax for Arbitrary Label Trees2024
  2. 2Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model using 3D Whole-body CT Scans2024 · 2 citations
  3. 3MedSAM2: Segment Anything in 3D Medical Images and Videos2025 · 7 citations
  4. 4Leveraging Multi-Text Joint Prompts in SAM for Robust Medical Image Segmentation2025 · 2 citations
  5. 5Necessity and Impact of Specialization of Large Foundation Model for Medical Segmentation Tasks2024