Abstract Rationale While machine learning (ML) models excel at predicting outcomes, their “black box” nature limits their use in clinical counseling. We developed an interpretable framework to deconstruct ML predictions into patient-specific risk factors, facilitating personalized risk communication for lung transplant candidates. Methods An XGBoost-Cox proportional hazards model was trained on the United Network for Organ Sharing (UNOS) registry, including 24,142 adult, primary lung transplant recipients. The model’s discrimination for 5-year survival was compared to the Composite Allocation Score (CAS) using time-dependent AUC. To illustrate the framework’s explanatory potential in a clinically intuitive scenario, we performed a demonstrative case study of two recipients with nearly identical CAS-derived survival estimates (∼53%) but highly divergent outcomes, using SHapley Additive exPlanations (SHAP) to analyze the model’s predictions. Results While our model achieved a statistically significant but modest improvement in discrimination over CAS (AUC: 0.653 vs. 0.598; p 0.001), its primary value lies in its interpretability. The case study involved a 68-year-old with Idiopathic Pulmonary Fibrosis (IPF) who died post-transplant and a 64-year-old with COPD who survived over nine years. Despite similar CAS profiles, our model assigned a markedly higher risk to the IPF patient (log-hazard: 0.59 vs. -0.00). SHAP analysis revealed that this risk was driven by the patient’s advanced age and diabetes duration, compounding their IPF diagnosis. Conversely, for the long-term survivor, the analysis highlighted the absence of a bilirubin measurement (a likely proxy for a more stable outpatient profile that does not require this test) as a key protective factor that counteracted the high risk posed by a 12.5-hour ischemic time. Conclusion Interpretable ML frameworks can translate complex model outputs into clinically intuitive, patient-specific narratives. By elucidating the rationale for a prediction, this approach provides a powerful tool to enhance patient education and manage expectations. Furthermore, it enables clinicians to critically appraise a model’s reasoning, including its reliance on proxy variables, thereby ensuring safer and more responsible deployment in clinical practice. This abstract is funded by: None
Alcudia et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: