To accurately evaluate flood susceptibility in Shenzhen and support long-term flood control planning, this study develops a GIS-based multi-model machine learning framework. Nine factors—including elevation, slope, and distance to rivers—were selected, with multicollinearity ruled out via Pearson correlation and VIF tests. A balanced sample set comprising 741 historical waterlogging points (2020–2024) and equal non-waterlogging sites was constructed. In addition to comparing five base models (Decision Tree, SVM, Logistic Regression, Naïve Bayes, LDA), the study introduces a voting ensemble for model integration and applies SHAP for both global and local interpretability. Key findings include: (1) improved predictive accuracy and robustness via ensemble learning (AUC = 0.8131), outperforming individual models; (2) flood susceptibility mapping reveals a distinct spatial pattern—higher risk in western coastal areas and lower risk in eastern mountainous zones—with 68.3% of historical waterlogging points located in high-susceptibility zones. The model is trained on waterlogging records from 2020 to 2024, which may not fully capture longer-term climatic or urban dynamics. This work directly supports sustainable urban development by providing a replicable framework for flood risk mitigation that reduces long-term economic and social vulnerabilities.
Yuan et al. (Fri,) studied this question.