e15517 Background: Recurrence following curative-intent hepatic resection for colorectal cancer liver metastases (CRLM) remains common, affecting up to 70% of patients mostly within the first postoperative year. Conventional clinicopathologic risk stratification tools offer limited discrimination and insufficiently capture biological and spatial heterogeneity. Artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches enable integration of high dimensional imaging, molecular, immune, and clinical data, and may improve postoperative recurrence risk stratification. Methods: We conducted a systematic review and random-effects meta-analysis comparing the discriminative performance of classical logistic regression (LR) and ML/DL/AI models for predicting postresection recurrence after curative-intent resection of CRLM. PubMed, Scopus, and Web of Science were searched from inception through January 2026 for studies reporting LR- or ML/DL/AI-based prediction of recurrence or time to recurrence, quantified using the area under the receiver operating characteristic curve (AUC) or concordance index. Eligible studies enrolled adults after curative resection and incorporated clinical, pathological, radiomic, or multi-omic predictors; studies lacking discrimination metrics were excluded. To avoid duplication, the best-performing model per study was selected. Discrimination estimates were logit-transformed and pooled using restricted maximum likelihood random-effects models, generating mean AUCs with 95% confidence intervals (CIs) and 95% prediction intervals (PIs). Funnel plot asymmetry was assessed using unweighted Egger tests. Results: Seven LR model evaluations were included, yielding a pooled AUC of 0.711 (95% CI 0.667-0.750; 95% PI 0.583-0.812), consistent with moderate discrimination for postresection recurrence. Seven ML/DL/AI model evaluations demonstrated a pooled AUC of 0.720 (95% CI 0.658-0.775; 95% PI 0.566-0.835), indicating comparable overall performance with a modestly higher point estimate. Egger testing showed no significant asymmetry for LR models (t = 1.65, df = 5, p = 0.160) while ML/DL/AI models exhibited significant asymmetry (t = 2.923, df = 5, p = 0.033), raising concern for potential publication or reporting bias. Conclusions: AI-, ML-, and DL-based models provide clinically meaningful discrimination for predicting recurrence risk following curative resection of CRLM, with pooled performance exceeding conventional risk stratification. Models incorporating imaging, molecular, and clinical features, may enhance risk stratification. Prospective validation, standardized recurrence endpoints, and external calibration are essential to support clinical implementation and postoperative risk stratification.
Seeburun et al. (Thu,) studied this question.