PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 19, 20250 citationsOpen Access

Large Language Models Improve Cancer Survival Prediction Using Real-World Clinical Notes

View Full Paper
NKNiklas KiermeyerTLTim LenfersADAmin Dada

Key Points

  • Incorporating LLM-derived features enhanced survival prediction significantly compared to traditional methods.
  • The C-Index for NSCLC improved from 0.64 to 0.72 with LLM integration, demonstrating greater predictive accuracy.
  • Using real-world clinical notes from over 3,500 patients, LLMs extracted key prognostic indicators effectively.
  • The study reveals that LLMs can reclassify risk in over 60% of patients, highlighting their potential clinical utility.

Abstract

In medical documentation, vast amounts of unstructured text are generated that are still underutilized in current prognostic models. We investigate the potential of self-hosted large language models (LLM) to extract clinically meaningful, patient-specific information from routine clinical notes for personalized risk stratification in cancer care. We collected real-world medical notes from 2,708 non-small cell lung cancer (NSCLC) patients and 814 colon cancer patients documented before treatment at a large comprehensive cancer center. LLMs extracted key prognostic indicators, including comorbidities, metastatic sites, and qualitative descriptors of patient condition, in a zero-shot manner without prior task-specific training. Integrating these LLM-derived features into machine learning models significantly improved the prediction of overall survival compared to TNM staging alone (C-Index: NSCLC, 0.72 vs 0.64; colon cancer, 0.70 vs 0.59), and surpassed models using text embeddings. Based on the LLM-informed risk scores, patients were stratified into four distinct risk groups, enabling reclassification of 61.4% of NSCLC and 68.3% of colon cancer patients. Analysis of model drivers revealed that LLM-derived factors, such as the physical condition, substantially modulated the prognostic impact of TNM stage. These findings highlight the potential of self-hosted LLM to extract clinically meaningful information from unstructured clinical documentation and support clinical decision-making.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kiermeyer et al. (2025) studied this question.

synapsesocial.com/papers/68af494dad7bf08b1ead4d7chttps://doi.org/10.1101/2025.08.17.25333835
Ask AI
Helpful
Bookmark
Share
View Full Paper