PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026BMC Medical Informatics and Decision Making0 citationsOpen Access

AI-assisted tumor board decision-making in pancreatic oncology

MMMarkus MergenFBFelix BuschBSBenjamin Schwarberg

Key Points

  • The research aims to assess the capability of an AI model to predict treatment decisions in pancreatic oncology tumor boards.
  • Evaluated LLaMA 3.3 (70b) model for decision-making in pancreatic cancer treatment.
  • Collected clinical documentation from 42 first-diagnosed cases in a real-world tumor board.
  • Applied different prompting strategies: zero-shot, advanced zero-shot, Chain-of-Thought, and few-shot prompting.
  • Measured performance with accuracy, micro- and macro-averaged F1 scores, and category-specific recall.
  • Achieved a highest overall accuracy of 78.6% using advanced zero-shot and CoT strategies.
  • High accuracy mainly from correctly classifying surgical and palliative cases, failing to identify any neoadjuvant therapy candidates.
  • Few-shot prompting identified NEO cases but reduced overall accuracy to 56.7%.
  • Indicates discrepancies in decision-making based on treatment complexity, showing potential risks in clinical applications.

Abstract

Abstract Background Pancreatic cancer requires nuanced, multidisciplinary treatment planning typically conducted within tumor boards. While Large Language Models (LLMs) have shown capabilities in medical reasoning, their ability to approximate complex, integrative decision-making in oncology remains underexplored. Methods This study evaluated the performance of LLaMA 3.3 (70b) in predicting tumor board decisions for newly diagnosed pancreatic cancer patients. Clinical documentation (including free-text imaging reports, pathology findings, and patient history) from 42 first-diagnosis cases discussed in a real-world tumor board was collected. The model was tasked with predicting one of three treatment options: surgical resection (SURG), neoadjuvant chemotherapy (NEO), or palliative therapy (PALL). Four prompting strategies were evaluated: zero-shot, advanced (adv.) zero-shot, Chain-of-Thought (CoT), and few-shot prompting. Performance was assessed using accuracy, micro- and macro-averaged F1 scores, and category-specific recall. Results The advanced zero-shot and CoT strategies achieved the highest overall accuracy of 78.6% and a micro-averaged F1 score of 0.786. However, this performance was driven primarily by the correct classification of majority classes (SURG and PALL). Crucially, both high-accuracy strategies failed to identify any of the neoadjuvant therapy candidates (Recall NEO = 0.00; 0/7 cases), systematically misclassifying them as palliative or surgical. While few-shot prompting improved the detection of neoadjuvant cases (Recall NEO = 1.00), it introduced substantial noise, reducing overall accuracy to 56.7%. LLaMA 3.3 (70b) demonstrates high concordance with tumor board decisions for clear-cut surgical or palliative cases but exhibits a critical systematic failure in identifying candidates for neoadjuvant therapy. The high global accuracy masks a significant safety limitation regarding the recognition of complex, intermediate-stage patients. Conclusion These findings suggest that current LLMs may approximate majority-class decisions but risk overlooking curative treatment pathways in nuanced scenarios, necessitating rigorous oversight and specific adaptation before clinical consideration

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mergen et al. (2026) studied this question.

synapsesocial.com/papers/69be36d46e48c4981c675febhttps://doi.org/10.1186/s12911-026-03444-x
Ask AI
Helpful
Bookmark
Share
View Full Paper