Key points are not available for this paper at this time.
This study investigates the determinants of driver fault attribution in road traffic crashes using a nationwide police database from Poland covering 2015–2024. The analysis focuses exclusively on drivers and applies three ensemble classifiers—Random Forest, Gradient Boosting, and Extreme Gradient Boosting (XGBoost)—to predict the likelihood of being designated as at fault. Group-wise cross-validation ensured methodological rigor by preventing information leakage between participants of the same crash. The tuned XGBoost model achieved a mean ROC AUC of 0.66, showing stable discrimination across folds. Feature-importance analysis identified possession of a valid driving licence and driving under the influence as the most influential predictors, followed by age, driving experience, and vehicle type. SHapley Additive exPlanations (SHAP) provided local interpretability, confirming that the global ranking of key variables was consistent at the individual prediction level. The results demonstrate that legal and behavioural determinants dominate fault attribution, while technical and demographic features exert smaller effects. The study highlights that improving data completeness—particularly on distraction and fatigue—offers greater potential for predictive gains than further algorithmic refinement. These findings contribute to an interpretable, data-driven understanding of crash responsibility, supporting evidence-based road-safety policy and enforcement strategies.
Artur Budzyński (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: