Background/Objectives: We developed and validated a multimodal magnetic resonance imaging (MRI) framework combining deep learning segmentation with radiomics to predict local recurrence in nasopharyngeal carcinoma (NPC). Methods: This retrospective two-center study included 1074 NPC patients treated between 2015 and 2019. Center 1 cases were split 8:2 into training and internal test sets, while Center 2 served for external validation. A multimodal Swin UNet model automatically segmented tumors from pretreatment T1-weighted, T2-weighted, and contrast-enhanced T1 (CET1) images. Radiomics features were extracted from expert-reviewed regions of interest, selected, and modeled using extreme gradient boosting for recurrence prediction. Results: The multimodal segmentation model maintained consistent but moderate Dice similarity coefficients (0.737, 0.666, and 0.726 for T1WI, T2WI, and CET1 in external validation). These values reflect the moderate overlap typical for nasopharyngeal carcinoma, given its highly infiltrative growth and ill-defined boundaries along complex anatomic interfaces. For local recurrence prediction, single-modality models reached external AUCs between 0.754 and 0.781. Importantly, the multimodal fusion model demonstrated numerical improvement over single modalities in the external validation set (e.g., vs. T1WI, p = 0.141), achieving an AUC of 0.910, accuracy of 0.908, sensitivity of 0.805, specificity of 0.946, and F1-score of 0.825. Conclusions: The multimodal MRI radiomics model, developed alongside a deep learning segmentation module, demonstrated favorable multicenter performance for evaluating NPC recurrence risk. The primary prognostic analysis was based on expert-reviewed regions of interest; a supplementary analysis using fully automatic segmentation masks yielded comparable, non-significantly different performance across all cohorts (Training AUC: 0.887; Internal Test AUC: 0.892; External Validation AUC: 0.885 vs. 0.910, p = 0.145), supporting the feasibility of future end-to-end deployment. Fusing multimodal features yielded numerical improvements over single-sequence models in external validation, providing a basis for post-treatment surveillance planning.
Yao et al. (2026) studied this question.