Sepsis, a life-threatening condition causing significant global mortality, requires rapid diagnosis and intervention. Although recent advances in machine learning have supported clinical decision-making, existing sepsis classification approaches exhibit several limitations, including inadequate temporal modeling of disease progression, lack of systematic hyperparameter optimization, fragmented interpretability approaches that do not fully address multi-stakeholder clinical needs, and challenges in achieving balanced sensitivity–specificity trade-offs. These limitations restrict effective extraction of knowledge from complex temporal clinical data and hinder actionable decision-making. To address these challenges, this work proposes BayeStack, a temporal knowledge-extraction framework that integrates Bayesian optimization-driven ensemble learning with hierarchical interpretability to optimize sepsis classification. This framework captures the progression of sepsis through multi-window temporal aggregation, performs optimal classification by applying AUROC-maximizing hyperparameter space exploration, and enables comprehensive clinical knowledge extraction by applying a three-level interpretability framework that includes global feature importance, population-level partial dependence analysis, and patient-specific contribution-level analysis. Evaluation results indicated that BayeStack achieved an AUROC of 0.99 with balanced sensitivity and specificity of 0.97, substantially outperforming all baseline methods (p<0.001). Ablation studies validated that temporal aggregation and data balancing contributed to performance improvements. A strong Spearman correlation (ρ=0.856) validated the feature ranking convergence and effectiveness of the ensemble strategy. The interpretability framework provides insights into complementary model behavior and extracts evidence-based clinical thresholds for priority-based treatment monitoring, thereby enabling robust clinical decision support. This first phase systematic integration framework of traditional machine learning models establishes baseline performance and explainability standards for subsequent deep learning advancements.
Geetha et al. (Mon,) studied this question.