PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 13, 2026ESMO Real World Data and Digital Oncology1 citationsOpen Access

Prompting large language models and evaluating inter- and intra-rater agreement for cancer progression assessment from radiology reports

TKT.Ó. KristjánssonAHA.F. HenriksenMHM.L. Hansen

Key Points

  • This study aims to evaluate large language models for annotating cancer progression in radiology reports and compare their performance with human annotators.
  • Analyzed 376 radiology reports from 184 patients with metastatic breast cancer.
  • Six human annotators classified reports as progressive disease or non-progressive disease.
  • Utilized a 'reverse questioning' strategy to assess five LLM models.
  • Bootstrapping estimated confidence intervals for non-inferiority comparisons.
  • Examined human intra-rater agreement for consistency.
  • The LLM framework achieved a mean Cohen's kappa of 0.82, indicating non-inferiority to human agreement (0.79).
  • Llama 3.1 model demonstrated 100% sensitivity, 90% specificity, and 84% F1 score.
  • Mean intra-rater variability among human annotators was 0.87.

Abstract

Background: Manual annotation of free-text radiology reports is time-consuming and costly, delaying real-world evidence (RWE) studies in oncology.This study aimed to evaluate the performance of large language models (LLMs) in annotating cancer progression from Danish free-text radiology reports.The objectives were to determine whether human-to-LLM inter-rater agreement was non-inferior to human-to-human agreement, establish human intra-rater agreement, and develop a framework for tuning LLM performance to RWE needs. Materials and methods:We identified 376 radiology reports from 184 patients with metastatic breast cancer from Danish electronic health records.Six human annotators, including two experts, classified radiology reports as progressive disease (PD) or non-PD.A 'reverse questioning' strategy was used to evaluate five LLM model series (Mistral, Gemma, Gemma 2, Llama 3, and Llama 3.1).Bootstrapping estimated confidence intervals (CIs) and assessed non-inferiority of the best-performing LLM ensemble compared with human agreement, using a non-inferiority margin of 0.1. Results:The LLM framework was non-inferior to human annotators with a mean Cohen's kappa of 0.82 (95% CI 0.74-0.89)for human-to-LLM versus 0.79 (95% CI 0.71-0.86)for human-to-human agreement (P < 0.001).The bestperforming ensemble model, Llama 3.1:70B, achieved 100% sensitivity, a specificity of 90%, and an F1 score of 84% on the test set.The mean human intra-rater variability was 0.87.Conclusions: The proposed LLM framework was non-inferior to human annotators in classifying cancer progression from free-text radiology reports.This offers significant potential for using LLMs as a tool for identifying tumor progression events in clinical assessment and research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kristjánsson et al. (2026) studied this question.

synapsesocial.com/papers/69b3acd302a1e69014cced46https://doi.org/10.1016/j.esmorw.2026.100689
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Toward expert-level medical question answering with large language models2025 · 908 citations
  2. 2Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models2023 · 3,806 citations
  3. 3Evaluating the Performance of ChatGPT in Ophthalmology2023 · 507 citations
  4. 4ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns2023 · 2,892 citations
  5. 5ChatGPT in radiology: A systematic review of performance, pitfalls, and future perspectives2024 · 124 citations