PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 23, 2026European Journal of Nuclear Medicine and Molecular Imaging2 citationsOpen Access

LLM-powered prostate cancer staging from PSMA-PET/CT reports using PROMISE v2

View Full Paper
DSDaniel SpitzlMMMarkus MergenLELukas Endroes

Key Points

  • This research aims to explore the use of large language models for automating prostate cancer staging from PSMA PET/CT reports.
  • Retrospective analysis of 1,696 PSMA PET/CT reports
  • Evaluated LLMs using four prompting strategies
  • Measured performance using accuracy, precision, recall, and F1 scores
  • Advanced Zero-shot prompting achieved micro-F1 of 0.65 for local staging
  • CoT prompting yielded micro-F1 of 0.79 for nodal and 0.84 for metastatic staging
  • Zero-shot approaches collapsed predictions into central categories, while advanced strategies were more balanced

Abstract

Accurate staging of prostate cancer is essential for guiding therapy and predicting outcomes. Prostate-specific membrane antigen (PSMA) PET/CT has become an established modality for disease assessment, and the PROMISE v2 framework provides standardized criteria for molecular imaging–based TNM (miTNM, molecular imaging TNM) classification. This study investigates the potential of large language models (LLMs) to automatically extract PROMISE v2 staging information from PSMA PET/CT reports. We retrospectively analyzed 1, 696 reports from first-diagnosis prostate cancer patients using the open-source Meta-Llama-3. 1-8B-Instruct (in the Q8₀ GGUF quantization) model, deployed securely within institutional infrastructure. Four prompting strategies were systematically compared: Zero-shot, advanced Zero-shot, few-shot, and chain-of-thought (CoT). Performance was evaluated using accuracy, precision, recall, and micro- and macro-averaged F1 scores. Advanced Zero-shot prompting achieved the highest performance for local staging (micro-F1 = 0. 65), while CoT prompting was superior for nodal (micro-F1 = 0. 79) and metastatic staging (micro-F1 = 0. 84). Error analyses revealed that Zero-shot approaches tended to collapse predictions into central categories, whereas advanced Zero-shot and CoT prompting yielded more balanced and stage-specific outputs. These findings demonstrate that LLMs can reliably map narrative PET/CT reports onto structured PROMISE v2 staging categories, with prompting strategy strongly influencing performance across staging dimensions. Our results highlight the feasibility of LLM-based workflows for standardizing prostate cancer staging, enabling reproducibility, and supporting large-scale outcome analyses.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Spitzl et al. (2026) studied this question.

synapsesocial.com/papers/69c08bcaa48f6b84677f980chttps://doi.org/10.1007/s00259-026-07847-w
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Testing and Evaluation of Health Care Applications of Large Language Models2024 · 577 citations
  2. 2Prostate Cancer Molecular Imaging Standardized Evaluation (PROMISE): Proposed miTNM Classification for the Interpretation of PSMA-Ligand PET/CT2017 · 634 citations
  3. 32022 Update on Prostate Cancer Epidemiology and Risk Factors—A Systematic Review2023 · 847 citations
  4. 4EAU-EANM-ESTRO-ESUR-ISUP-SIOG Guidelines on Prostate Cancer—2024 Update. Part I: Screening, Diagnosis, and Local Treatment with Curative Intent2024 · 1,253 citations
  5. 5Exploring Multilingual Large Language Models for Enhanced TNM Classification of Radiology Report in Lung Cancer Staging2024 · 20 citations