Key points are not available for this paper at this time.
Background: Predicting rehabilitation outcomes at admission supports tailored therapy plans and efficient use of resources for patients undergoing intensive inpatient rehabilitation, including those with stroke, orthopedic, and other neurological conditions. Nonetheless, current machine learning (ML) methods face limitations, including the ceiling effect in absolute functional gain measures, the uniform treatment of diverse patient groups, and reliance on black-box models that lack clinical transparency. Methods: This retrospective observational study analyzed a fully anonymized, publicly available dataset of 3419 patients admitted to the intensive rehabilitation unit at IRCCS San Raffaele Hospital, Rome, Italy, from 2015 to 2018. To mitigate the ceiling effect, a normalized Barthel Index gain metric was developed. K-means clustering (K = 2, trained solely on the training set) identified patient admission profiles based on functionality, which were then used as predictive features. Eight machine learning classifiers were tested across three groups (All Patients, Orthopedic, Neurological). SHAP-based explainability was employed at four levels: global, diagnostic group, patient functional profile, and individual. Finally, clinical decision rules and bedside stratification profiles were derived and validated with an internal held-out test set (n = 684). Results: Normalization significantly increased the correlation between admission BI and gain (r = 0.130 to r = 0.520), supporting the presence of a ceiling-related limitation in absolute gain metrics. Two distinct functional admission profiles with statistically significant group differences were identified—High-Burden (38% below-median recovery) and Moderate-Burden (21%)—with cluster membership the third most important predictor (13.9% SHAP importance). The highest AUC-ROC values were 0.831 for all patients (XGBoost), 0.864 for neurological patients (Gradient Boosting), and 0.839 for orthopedic patients (Gradient Boosting). Multilevel SHAP analysis showed age as the primary predictor for neurological patients (mean |SHAP| = 0.360) but the third for orthopedic patients (0.350), highlighting clinical relevance. Validation using SHAP values from the Gradient Boosting model showed a Spearman correlation of ρ = 0.925 (p = 1.13 × 10−30), with eight of the top ten features overlapping, indicating that these patterns are not model-specific but reflect the underlying data. Risk zone stratification found 80.7% of patients in high-confidence zones (accuracy > 80%). The clinical decision rules achieved 70.8% accuracy with full transparency, and the elderly (≥75 years) combined with a low BI (<25) profile showed an 89.6% model accuracy with only 10.4% recovery above the median. Conclusions: This explainable, profile-informed ML pipeline addresses key methodological limitations in predicting rehabilitation outcomes. It also provides a foundation for integrating models into clinical practice, pending prospective, external validation of the results. Before clinical implementation, validation across multicenter cohorts is essential.
Hawamdeh et al. (Sun,) studied this question.