PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 16, 2025Traffic Injury Prevention2 citations

Developing an XGBoost based model to predict the probability of truck crashes driven by macro operation and insurance data

View Full Paper
YWYiping WuHZHongpeng ZhangPSPeng Song

Key Points

  • The predictive model achieved an accuracy of 87.59%, showcasing its effectiveness in identifying truck accident probabilities.
  • Key metrics, including a recall rate of 84.21% and an F1 score of 85.33%, highlight the model's performance in real-world applications.
  • Dimensionality reduction analysis identified minimal input requirements, emphasizing the model's efficiency in handling operational data.
  • SHAP values and PCA analysis revealed that roadway familiarity and segment type significantly influence the likelihood of accidents.

Abstract

Truck accidents caused significant casualties usually. Establishing a scientific truck accident prediction model and identifying the primary causes are crucial for proactive accident prevention. The proposed model was developed using annual operational behavior data and corresponding insurance claim information from commercial trucks. Prior to model training, multicollinearity among predictor variables was addressed to ensure model interpretability and stability. Model performance was evaluated using recall, F1 score, and overall prediction accuracy, including external validation with a temporally separated dataset from the same driver population. To reduce input data dependency, an input dimensionality reduction analysis was conducted to determine the minimal data requirements. SHAP (Shapley Additive Explanations) values and principal component coefficients were employed to extract the main factors influencing truck accidents. The truck accident prediction model with a recall rate of 84.21% and an F1 score of 85.33%. The prediction accuracy of our developed model reached 87.59% when using new data from the same group of truckers in the subsequent year for validation. Additionally, the minimum data requirement set for our developed model was found to be the feature combination of load capacity, road segment type, and driving time, through analyzing the relationship between model prediction accuracy and feature inputs with different dimensions. Based on the suggested model inputs, the recall rate and F1 score of the prediction model are 86.84% and 84.62%, respectively. The main influencing factors analyzed by SHAP values and Principal Component Analysis (PCA) coefficients indicated that the trucker's familiarity with the road and the type of road segment significantly impact the probability of accident occurrence. This research innovatively establishes a macro data-driven truck accident prediction model alleviating the difficulty of data collection as well as guaranteeing the prediction accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68d453a431b076d99fa5997dhttps://doi.org/10.1080/15389588.2025.2545002
Ask AI
Helpful
Bookmark
Share
View Full Paper