PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 14, 2026Scientific Reports1 citationsOpen Access

Performance of GPT-based large language models in hepatocellular carcinoma stratification: liver function assessment, BCLC staging, and treatment recommendations

MMMax MasthoffAZAmelie ZipserMPMichael Praktiknjo

Key Points

  • This research aims to assess the efficacy of GPT-based large language models in analyzing data to guide clinical decisions for patients with hepatocellular carcinoma.
  • Analyzed clinical, radiological, and laboratory data from 106 patients with HCC.
  • Compared outputs from four GPT versions against expert consensus and tumor board decisions.
  • Evaluated time and cost savings of GPT versus clinical staff.
  • GPT versions achieved over 85% accuracy in liver function assessments, with MELD calculations prone to more errors.
  • BCLC staging accuracy varied from 46.2% for version 4 to 84.0% for version o3, with main errors from misclassified radiological reports.
  • Reasoning-optimized models (o1, o3) reached treatment recommendation accuracy of 90.6%, and in 9-14% of cases, GPT was more guideline-concordant than tumor board decisions.

Abstract

Abstract Large language models (LLMs) like GPT have been proposed to support complex clinical decision-making. This study evaluated the performance of GPT-based LLM in analyzing clinical, radiological, and laboratory data from patients with hepatocellular carcinoma (HCC) to assess liver function, assign BCLC stage, and recommend treatment. Data from 106 HCC patients (82% male, median age 65 22–86) were compiled into anonymized integrated reports. Four GPT-versions (4, o1, o3, 5.4) were prompted—using both short and long instructions—to calculate MELD, ALBI, and Child–Pugh scores, assign BCLC stage, and generate treatment recommendations based on current guidelines. Outputs were compared to expert consensus and tumor board decisions. Errors were categorized by type and source. Time and cost analyses compared GPT to clinical staff. All GPT versions achieved high accuracy (> 85%) in liver function assessment, with MELD calculation being the most error-prone. BCLC staging accuracy ranged from 46.2% (version 4) to 84.0% (o3), with misclassification of radiological reports as the main error source. Reasoning-optimized models (o1, o3) performed best for treatment recommendations, achieving an overall accuracy (correct suggestions and acceptable alternatives) of up to 90.6%. In 9–14% of cases, GPT suggestions were retrospectively more guideline-concordant than tumor board decisions. GPT processing was significantly faster and reduced costs by approximately 300- to 1300-fold compared to clinical staff. GPT-based LLMs show potential as decision-support tools for liver function assessment, BCLC staging, and treatment guidance in HCC. Particularly with reasoning-optimized models and detailed prompting, LLMs may serve as valuable adjuncts in multidisciplinary HCC workflows. However, a non-negligible error rate requires expert oversight and further model refinement.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Masthoff et al. (2026) studied this question.

synapsesocial.com/papers/6a2e47cdb1cc60ccdea8c3b5https://doi.org/10.1038/s41598-026-56992-7
Ask AI
Helpful
Bookmark
Share
View Full Paper