PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026IEEE Journal of Biomedical and Health Informatics0 citations

RT-SAM: Visual-Prompt Fusion and Uncertainty Enhancement for Nasopharyngeal Carcinoma Radiotherapy Target Delineation

View Full Paper
HKHee Guan KhorXYXin YangYSYihua Sun

Key Points

  • The aim is to develop RT-SAM for automated delineation of clinical target volumes in nasopharyngeal carcinoma radiotherapy.
  • Adapted Medical Segment Anything Model 2 (MedSAM-2) and 2D U-Net for contouring
  • Generated multi-modal prompts from specialist network predictions
  • Implemented Visual-Prompt Fusion Attention (ViPFA) for enhanced feature interaction
  • Applied Uncertainty-Enhanced Prediction Adjustment (UEPA) for model refinement
  • Achieved a mean DICE coefficient of 0.796 ± 0.033
  • RT-SAM contours were clinically indistinguishable from expert delineations
  • 75% of RT-SAM contours received superior quality ratings compared to manual expert contours
  • Clinically acceptable ratings in over 97% of assessed cases

Abstract

Precise delineation of the clinical target volume (CTV) and nodal CTV (CTV₍₃) is crucial for effective radiotherapy planning in nasopharyngeal carcinoma (NPC). Manual contouring is labor-intensive and subject to substantial inter-observer variability, particularly in regions with complex anatomy and indistinct boundaries. This study presents RT-SAM, a novel framework that adapts the Medical Segment Anything Model 2 (MedSAM-2) for automated CTV (i. e. , primary CTV and CTV₍₃) contouring in NPC computed tomography (CT) images. The framework synergistically integrates a generalist foundation model (MedSAM-2) with a domain-specific specialist network (2D U-Net) through three principal contributions: (1) automated generation of multi-modal prompts-comprising mask, bounding box, and point representations-derived from specialist network predictions to guide the generalist model; (2) a Visual-Prompt Fusion Attention (ViPFA) mechanism that optimizes feature-prompt interactions through bidirectional cross-modal attention; and (3) an Uncertainty-Enhanced Prediction Adjustment (UEPA) mechanism that enhances model robustness via confidence-based refinement and selective domain adaptation. Comprehensive evaluation on a multi-center cohort of 256 clinical NPC cases from Sun Yat-sen University Cancer Center and 212 public NPC cases from the SegRap2025 lymph node CTV dataset using 5- fold cross-validation demonstrates that RT-SAM achieves a mean DICE coefficient of 0. 796 0. 033 (mean standard deviation), significantly outperforming current state-of-the-art methods. Clinical validation by eight radiation oncologists demonstrates that RT-SAM contours are clinically indistinguishable from expert delineations in blinded Turing assessments, achieve superior quality ratings in 75% of comparisons with mean scores of 2. 73 for RT-SAM versus 2. 66 for manual expert contours, and attain clinically acceptable ratings in over 97% of cases. These results demonstrate that RT-SAM is a clinically feasible solution for automated CTV contouring, with strong potential to standardize treatment planning and mitigate inter-observer variability in NPC radiotherapy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khor et al. (2026) studied this question.

synapsesocial.com/papers/69aa6f0d531e4c4a9ff5930ahttps://doi.org/10.1109/jbhi.2026.3669979
Ask AI
Helpful
Bookmark
Share
View Full Paper