PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 17, 2026Future Transportation2 citationsOpen Access

Prediction of Large-Scale Traffic Accident Severity in Qatar: A Binary Reformulation Approach for Extreme Class Imbalance with Interpretable AI

View Full Paper
MAMohammed AlshriemYYYin Yang

Key Points

  • The aim is to develop machine learning frameworks to predict traffic accident severity while addressing class imbalance and ensuring interpretability.
  • Systematic preprocessing of 588,023 accident records from a national dataset
  • Comparative analysis of multi-class, binary, and cascaded two-stage approaches
  • Evaluation of six classifiers across different encoding and balancing strategies
  • Hyperparameter tuning with 5-fold stratified cross-validation
  • Temporal validation using data from 2020-2025
  • The binary LightGBM classifier achieved a balanced accuracy of 71.04% and an AUC-ROC of 0.772
  • Sensitivity was found to be 61.03% while specificity reached 81.05%
  • Temporal analysis indicated that the time period was the most significant predictor of accident severity
  • The findings support targeted enforcement strategies for improving pedestrian safety.

Abstract

Road traffic injuries represent one of the most critical public health challenges in the Gulf region. Predicting traffic accident severity is therefore a critical component of evidence-based road safety management. In this study, we develop machine learning frameworks for predicting traffic accident severity using Qatar’s national dataset (2020–2025), addressing extreme class imbalance and interpretability. A dataset of 588,023 accident records was systematically preprocessed from 1,000,500 raw reports. We compare three approaches: multi-class (four severity levels), binary (Safe vs. Severe), and cascaded two-stage (combining both). Six classifiers were evaluated across two encoding methods and three balancing strategies. Systematic hyperparameter tuning with 5-fold stratified cross-validation was performed for all models. The binary LightGBM classifier achieved BA = 71.04%, AUC-ROC = 0.772, Sensitivity = 61.03%, and Specificity = 81.05%, demonstrating superior performance over multi-class approaches. Temporal validation on 2025 data (trained on 2020–2024 data) supported good temporal generalization. Analysis of 10,000 test instances identified the time period as the dominant predictor of accident severity. The binary LightGBM framework provides an interpretable and effective approach for severe accident identification and risk prioritization, with SHAP findings supporting targeted temporal enforcement and pedestrian safety as evidence-based policy priorities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alshriem et al. (2026) studied this question.

synapsesocial.com/papers/69e1ce605cdc762e9d857777https://doi.org/10.3390/futuretransp6020088
Ask AI
Helpful
Bookmark
Share
View Full Paper