Objectives To develop and internally evaluate a probability-level stacked ensemble model integrating diffusion-weighted imaging (DWI) and apparent diffusion coefficient (ADC) magnetic resonance imaging (MRI)-based imaging predictions with routinely available clinical predictions to predict 90-day poor functional outcome in patients with acute ischemic stroke. Methods This retrospective single-center study included 562 patients with acute ischemic stroke who underwent DWI and ADC MRI within 1 to 7 days after symptom onset. Poor outcome was defined as 90-day modified Rankin Scale 2. Patients were assigned at the patient level to training ( n = 393), validation ( n = 84), and held-out internal test ( n = 85) sets. A ResNet50-based imaging model and a support vector regression (SVR)-based clinical model were integrated using logistic regression as a probability-level meta-learner. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, calibration, decision curve analysis, and the DeLong test. Results In the held-out internal test set, the fusion model achieved the highest numerical AUC (0.951, 95% CI: 0.908–0.994), with sensitivity of 0.889 and specificity of 0.911. The DeLong test showed that the fusion model significantly outperformed the imaging model ( p 0.05), but not the clinical model or Wouters 2018 model. Brier scores were 0.077 for the Wouters 2018 model, 0.089 for the clinical model, 0.094 for the fusion model, and 0.122 for the imaging model. Decision curve analysis showed that the fusion model had positive net benefit across threshold probabilities from 0.10 to 0.60 and exceeded the imaging model, but did not consistently exceed the clinical model or Wouters 2018 model. Conclusion The fusion model showed high internal discrimination and improved performance compared with imaging alone, but its incremental value over clinical models was limited in calibration and decision curve analyses. External validation, recalibration, and prospective evaluation are required before broader clinical use.
Ji et al. (Tue,) studied this question.