In this study, we use an ensemble learning approach for diabetes risk prediction using the Pima Indians Diabetes dataset.Missing values were handled using median imputation, and the feature "SkinThickness" was removed as it had very little impact on modelperformance. Random Forest, XGBoost, and LightGBM models were evaluated, and the final prediction was obtained using a soft votingensemble strategy. The proposed ensemble achieved ROC-AUC = 0.821, recall = 0.685, and F1 = 0.667. SHAP-based interpretation andstability analysis consistently identified Glucose, BMI, DiabetesPedigreeFunction, Age, and Pregnancies as dominant predictors. Riskstratification, fairness evaluation, and decision curve analysis demonstrated clinically interpretable and reliable behavior. The resultssuggest that explainable ensemble models provide effective support for early diabetes screening.
Aniket Mishra (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: