AbstractPurpose Interpretability is highly desirable for oncologic outcome prediction, as it increases the level of transparency and trustworthiness of the model. This model characteristic is particularly relevant in the setting of modest sample size. Existing work has focused on Shapley Additive Explanations to provide post-hoc explanations on black-box models. These models are not intrinsically interpretable. In this study, we investigated the applicability of an intrinsically interpretable glass-box model, Explainable Boosting Machine (EBM), for hypothesis generation from an early-stage lung cancer data set. Methods We applied EBM to a stripped dataset on post-radiotherapy lung cancer recurrence, aiming to extract as much information as possible by using pristine EBM configurations. We compared the key features ranked by EBM with those identified through univariate statistical analysis. Additionally, we benchmarked its performance against logistic regression (LR) and random forest (RF) models, while also evaluating the hypotheses generated by EBM. Results EBM identified primary tumor size and body mass index (BMI) as the most prognostic features, aligning with the results of the univariate analysis. Its interpretability provides safeguards against misinterpretation; the model revealed potential age-related bias in this single-arm dataset and possible confounding interactions between race and BMI. EBM yielded competitive performances and more interpretable insights compared with LR and RF but was not immune from generalizability challenges arising from limited data. Conclusions The modest performance prevents EBM from being used as clinical decision support tool, when applied to limited data. However, its interpretable, glass-box nature makes it useful for hypothesis generation.
Zhang et al. (Sun,) studied this question.