Abstract Objective To develop a comprehensive risk stratification model to identify patients at high risk of cesarean surgical site infection (cSSI) based on information available at the time of hospital discharge. Methods This is a retrospective cohort study based on datasets from two independent level IV academic medical centers to improve generalizability. The first study includes 2336 cesarean deliveries for analysis with 109 instances of cSSI (4.7%). The second study includes 2936 cesarean deliveries with 72 instances of cSSI (2.5%). The first study cohort was utilized for both model training and internal validation, using the 15 shared features present in both datasets. Random search with cross‐validation was used to choose the type of machine learning (ML) model and its hyperparameters, with area under the receiver operating characteristic curve (AUC) as the validation metric. The second study cohort was then used to evaluate testing accuracy (external validation) and generalization capability of this model. The predictive importance of each feature was evaluated using permutation importance applied to the second cohort. Risk stratification cutoff scores were defined based on centiles over the first patient cohort and evaluated on the second cohort. Results The ML model with the highest validation AUC was an XGBoost classifier. This trained model achieved an AUC of 0.78 on the training data and a 0.72 cross‐validated AUC. When tested on data from the second study cohort, it achieved an AUC of 0.69. The features showing the strongest predictive generalization, as measured by permutation importance, were length of hospital stay, primary indication for cesarean delivery, and estimated blood loss. Using the second cohort to test our model, the model shows a 22.7% positive predictive value (PPV) of cSSI among women in the highest risk group, and a 0.8% PPV among women in the bottom risk group, compared with 2.5% cSSI incidence for the study population at large. Conclusion We developed and validated a predictive ML model for cSSI with information available to healthcare professionals at the time of hospital discharge. This model suggests that the most predictive risk factors for cSSI include length of hospital stay, primary indication for cesarean delivery, and estimated blood loss. Additionally, the model provides reasonable risk stratifications for individual patients. This tool may be used for identifying subgroups at higher risk of cSSI and identifying those who may require closer surveillance or clinical intervention.
Lobel et al. (Thu,) studied this question.