River flooding represents the most prevalent and destructive type of natural disaster, exerting profound adverse impacts on human health, as well as on natural and built environments, and permanently altering land surface characteristics. This study aims to model river flood risk within the Sago–Sile River Catchment Southern Ethiopia using two machine learning algorithms (MLAs): Random Forest (RF) and Support Vector Machine (SVM). A total of 428 flood and non-flood event data points were compiled and partitioned into training (70%) and testing (30%) datasets. Nineteen flood conditioning factors were initially considered, and the Correlation Attribute Subset (CAS) selection method was employed to identify the most influential predictors. Model performance was evaluated and compared using both prediction and success rate analyses. Flood vulnerability was assessed based on two principal dimensions: (i) economic vulnerability, represented by land use and tropical livestock units (TLU), and (ii) social vulnerability, represented by population density, age, gender, and household head characteristics. Factor weights were determined through expert judgment using the Analytic Hierarchy Process (AHP). The spatial distribution of flood exposure classes was derived using the Frequency Ratio (FR) method. Fourteen key conditioning factors—Elevation, Distance to shoreline, Topographic Roughness Index (TRI), Road proximity, Height Above Nearest Drainage (HAND), Normalized Difference Vegetation Index (NDVI), Distance to river, Slope, Drainage density, Topographic Position Index (TPI), Curve Number (CN), Topographic Wetness Index (TWI), Precipitation, and Land cover—were retained for flood exposure mapping. The RF model demonstrated superior predictive capability, yielding high accuracy in both prediction and success phases (Accuracy = 0.906 and 0.970; Kappa = 0.81 and 0.94; RMSE = 0.236 and 0.1523; ROC-AUC = 0.986 and 0.997, respectively). The consistency index (CI) for social vulnerability criteria was 0.015, indicating reliable weighting consistency. According to the results, approximately 21,396 inhabitants, 12,831.2 TLU, 7,578.82 ha of farmland, 528.1 ha of built-up area, and 31.2 km of road infrastructure (including two bridges) are exposed to river flood hazards. The river flood risk map, classified using a quintile scheme, delineated five risk zones: very low (29%), low (26%), moderate (28%), high (10%), and very high (7%). Spatial analysis revealed that most flood-prone areas occur at elevations between 1094–1175 m (99%), within low-relief terrain (roughness = 0), and in close proximity to roads (within 1.3 km, 54.8%). Overall, the findings demonstrate that the integration of RF with CAS feature selection provides a strong framework for river flood risk mapping, particularly in data-scarce or ungauged catchments where flood behavior and environmental interactions are insufficiently understood.
Dana et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: