Crash data often suffer from extreme class imbalance with nonfatal crashes vastly outnumbering severe or fatal cases, making it difficult to model rare but critical outcomes. These challenges are further exacerbated in low- and middle-income countries due to poor data quality. This study examines the impact of class imbalance correction approaches on the stability and interpretability of crash severity models in Kathmandu Valley, Nepal, using police-reported data of 5 years. Three modeling approaches are compared: standard multinomial logit (MNL), MNL with the synthetic minority oversampling technique (SMOTE), and MNL with class weights. The study introduces two new metrics, the log-odds ratio distance (LORD) and variable rank agreement coefficient (VRAC), to assess the stability of coefficients and variable importance rankings across models. The results show that data-balancing approaches (G-mean≈0.23) significantly outperform the standard MNL model (G-mean=0.00) in identifying severe crash outcomes, demonstrating substantial improvement in balanced classification performance. Further, the MNL with class-weights model shows better stability, slightly improved performance, and more consistent effect estimates for determining key factors, with lower LORD (0.1007 versus 0.2013) and higher VRAC (0.7868 versus 0.7640) than MNL with SMOTE. Critical determinants of crash severity include temporal factors such as weekends, nighttime, and monsoon season; vehicle types such as trucks and motorcycles; crash reasons such as high speed, alcohol influence, and overtaking; and crash types involving pedestrians. The study contributes methodological innovations and context-specific insights by highlighting the need to address class imbalance and supports evidence-based decision-making in rapidly developing urban areas facing growing road safety challenges.
Devkota et al. (Wed,) studied this question.