PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 14, 2026JCO Clinical Cancer Informatics0 citations

Prompt Engineering for Eastern Cooperative Oncology Group Status Extraction: Comparing Large Language Model Techniques

View Full Paper
MDMeenakshi DubeyNational University of SingaporeKCKok Joon ChongNational University of SingaporeYPYuba Raj PunNational University of Singapore

Key Points

  • This research aims to evaluate different methods for extracting ECOG performance status from unstructured clinical notes using large language models.
  • Evaluated four approaches to extract ECOG status from clinical notes.
  • Used a rule-based natural language processing algorithm and various prompting techniques.
  • Compared performance based on binary and three-class outcomes using adapted evaluation metrics.
  • Applied techniques to notes from patients with non–small cell lung cancer, multiple myeloma, or ovarian cancer.
  • Chain-of-thought and double filtering achieved 94% accuracy, outperforming simpler methods.
  • Double filtering showed the highest specificity (0.91) and positive predictive value (0.93).
  • Chain-of-thought achieved the highest sensitivity (0.98).
  • Both advanced approaches received higher satisfaction ratings in human evaluations.
  • Results for ECOG ≥2 were imprecise due to a small sample size.

Abstract

PURPOSE Eastern Cooperative Oncology Group (ECOG) performance status is critical for cancer patient management, yet it is often documented only in unstructured clinical notes. This study compares several approaches to extract ECOG status from oncology notes, focusing on advanced prompting techniques for large language models (LLMs). METHODS We evaluated four ECOG extraction approaches on unstructured clinical notes from patients with non–small cell lung cancer, multiple myeloma, or ovarian cancer (2017-2021). The approaches were a rule-based natural language processing algorithm, simple LLM prompting, and two advanced prompts (chain-of-thought and Double Filtering) using a domain-tuned LLM (LLAMAv3.2). Performance was measured on a binary outcome (any ECOG documented v none) and a three-class outcome (ECOG 0-1 v ≥2 v none) and via an adapted QUEST questionnaire for human evaluation. RESULTS Both CoT and double filtering technique (DFT) achieved 94% accuracy, outperforming the rule-based method (91%) and simple prompting (86%). DFT had the highest specificity (0.91) and positive predictive value (PPV; 0.93), whereas CoT attained the highest sensitivity (0.98). In the QUEST evaluation, DFT and CoT scored higher on output quality, reasoning, bias reduction, and user satisfaction than the simple prompt. DFT received the top satisfaction rating. In the three-class analysis, DFT and CoT again performed best (accuracy 0.91 v 0.87) and DFT was most sensitive for ECOG ≥2 cases. Estimates for ECOG ≥2 remained imprecise because of the small sample (n = 20). All methods sometimes hallucinated ECOG status. CONCLUSION Advanced LLM prompting improved ECOG extraction over basic methods. DFT and CoT each showed specific strengths (DFT had higher PPV and user satisfaction; CoT achieved higher sensitivity). These approaches appear to be generalizable across cancer types. Key implementation considerations include computational cost and human oversight. Overall, advanced prompting can standardize ECOG documentation, accelerate patient cohort identification, and inform personalized treatment planning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dubey et al. (2026) studied this question.

synapsesocial.com/papers/699011712ccff479cfe58201https://doi.org/10.1200/cci-25-00226
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A framework for human evaluation of large language models in healthcare derived from literature review2024 · 356 citations
  2. 2A framework for evaluating clinical artificial intelligence systems without ground-truth annotations2024 · 24 citations
  3. 3Toxicity and response criteria of the Eastern Cooperative Oncology Group1982 · 11,785 citations
  4. 4A large language model for electronic health records2022 · 865 citations
  5. 5CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis2025 · 9 citations