Key result
Digital pathology foundation model predicts biomarkers and survival across 12 cancer types from H&E slides.
Why the study?
Machine learning in oncology has been limited by high-resolution imaging modalities and scarce supervised data, but self-supervised foundation models can bypass these constraints.
A digital pathology foundation model trained on H&E slides can accurately predict biomarkers and survival across multiple cancer types, including breast cancer, without requiring clinical data.
Does not support clinical adoption for biomarker or survival prediction; leaves open the role of pathology foundation models pending prospective validation.
Background: The application of machine learning methods to oncology has historically been challenging due to the high resolution of medical imaging modalities and scarcity of downstream supervised data. Over the past years, however, foundation models have enabled the field to bypass these constraints by leveraging self-supervised learning on large quantities of unsupervised imaging data. These foundation models learn effective latent representations for the various morphologies seen throughout the training data. These representations can then be used downstream for supervised machine learning tasks with minimal additional training. Methods: A digital pathology foundation model was trained using self-supervised learning method DINOv2 on 250M patches extracted from 260k whole slide images (WSIs). We evaluated the model on biomarker classification and survival analysis tasks across a total of 17 cohorts covering 12 cancer types. For all evaluations, we kept the foundation model frozen and trained regressors or classifiers on top of mean-pooled embeddings obtained by passing hematoxylin and eosin (H&E)-stained slide patches through the model. No clinical data was used for any evaluation. For all tasks, we perform 5-fold cross-validation and report AUROC for classification tasks and C-index for survival tasks. For classification tasks, we used k-NN (k=20), and for survival tasks, we used the Cox Proportional Hazards model. Results: Downstream models performed particularly well in breast cohorts, with mutation classification AUROCs of 0.70 ± 0.07 (TP53) and 0.71 ± 0.18 (CDH1) in CPTAC-BRCA, and a breast cancer-specific survival (BCSS) C-index of 0.67 ± 0.06 in PLCO-Breast. The models also predicted disease-specific survival (DSS) in other cancer types, including bladder (PLCO) with a C-index of 0.77 ± 0.09 despite only 39 events among 285 unique patients. AI models also predicted PIK3CA mutation status in the CPTAC-GBM cohort (AUROC of 0.73±0.17), TP53 in CPTAC-LUAD (0.63±0.10), and homologous recombination deficiency (HRD) in pan-cancer TCGA cohorts (SARC: 0.80±0.06; STAD: 0.68±0.06; UCEC: 0.67±0.07), and microsatellite instability (MSI) in TCGA-STAD (0.69±0.06). See Table 1 for results for all tasks and cohorts. Conclusions: Pathology foundation models are applicable to a wide variety of tasks across a range of cancer subtypes, even in spite of sparse data and naive model architectures. With the right data, similar methodology could be used to train predictors of recurrence risk, metastasis risk, treatment benefit, and more. Citation Format: J. Cappadona, J. Witowski, K. Zeng, J. Park, B. Machura, K. Geras. Pan-cancer ai foundation models yield accurate biomarker and survival predictions in breast cancer [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2025; 2025 Dec 9-12; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(4 Suppl):Abstract nr PS3-06-10.
No takes yet. Share an insight, caveat, or question.
Cappadona et al. (2026) studied Pan-cancer, including breast cancer. Digital pathology foundation model (DINOv2) was evaluated on Biomarker classification (AUROC) and survival analysis (C-index). A digital pathology foundation model accurately predicted biomarkers and survival across 12 cancer types using only H&E slides, achieving a breast cancer-specific survival C-index of 0.67 ± 0.06.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: