Employee attrition represents a persistent organizational challenge with substantial financial and strategic consequences. Although machine learning techniques have increasingly been applied to attrition prediction, prior research has largely emphasized isolated models and overall accuracy, frequently overlooking severe class imbalance, pipeline-level design choices, and the translation of predictive outputs into operational and economic insight. This study develops a comprehensive, end-to-end machine learning framework for employee attrition prediction that systematically evaluates preprocessing and resampling strategies, establishes statistically grounded performance rankings, and integrates explainability with business-level interpretation. Predictive pipelines were constructed by combining standard and robust scaling with random undersampling, SMOTE, and ADASYN, and were implemented using logistic regression, random forest, XGBoost, and LightGBM models. All configurations were evaluated using stratified five-fold cross-validation and minority-sensitive measures including ROC–AUC, recall and MCC. The best-performing pipelines achieved ROC–AUC values up to 0.86, recall levels around 0.74, and MCC approaching 0.45, indicating strong discrimination together with substantially improved minority detection. Bayesian hyperparameter optimization further refined the selected configuration, and Friedman-based ranking identified logistic regression and LightGBM as the most stable model families. SHAP explainability revealed overtime exposure, compensation, satisfaction, and tenure-related variables as dominant drivers of attrition risk. Business value–driven threshold analysis indicated substantial potential cost savings, and risk profiling uncovered structurally distinct attrition-prone employee segments.
Sinap et al. (Mon,) studied this question.