PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 12, 2026Scientific Reports0 citationsOpen Access

In-vehicle 3D vision for perceiving dangerous driving behaviors

WLWuhuan LiChongqing Normal UniversityJLJun LuChongqing Normal UniversityKTKanlun TanState Key Laboratory of Vehicle NVH and Safety Technology

Key Points

  • This study aims to improve the detection of dangerous driving behaviors using a comprehensive 3D pose estimation framework.
  • Constructed a dual-view 3D pose dataset with ten typical driving behaviors using a Time-of-Flight camera.
  • Developed a lightweight end-to-end pipeline using an anchor-based regression model for 3D pose estimation.
  • Integrated pose estimation with a graph-based architecture for action recognition and real-time monitoring.
  • Achieved 96.02% accuracy in 3D pose estimation.
  • Secured 98.0% accuracy in behavior recognition under real-time conditions.
  • Demonstrated a computational cost of 1.49 G FLOPs with an inference latency of 0.0375 seconds per sample.

Abstract

Accurate identification of dangerous driving behaviors is critical for accident prevention and occupant protection. However, most existing in-vehicle driver monitoring systems rely primarily on facial or head motion analysis, which fails to capture full-body driving behaviors and raises privacy concerns due to dependence on RGB or near-infrared imaging. In addition, these systems often exhibit limited robustness under low-light conditions. To address these limitations, this study proposes a comprehensive depth-based framework for in-vehicle 3D human pose estimation and dangerous driving posture recognition. First, a large-scale dual-view 3D pose dataset encompassing ten typical driving behaviors is constructed using a Time-of-Flight (ToF) camera. Based on this dataset, we develop a lightweight end-to-end pipeline in which an anchor-based regression model estimates the 3D poses of 16 driver keypoints, followed by an enhanced ST-GCN++ architecture for skeleton-based action recognition. By integrating pose estimation with graph-based temporal modeling, the proposed method effectively distinguishes visually similar hazardous behaviors. To facilitate real-world deployment, the algorithm is further integrated into a software system that enables closed-loop pose monitoring and hierarchical intervention. Experimental results verify that the proposed method achieves 96.02% accuracy in 3D pose estimation and 98.0% accuracy in behavior recognition. With a computational cost of only 1.49 G FLOPs and an inference latency of 0.0375 s per sample, the system achieves real-time performance (27-28 FPS) on an automotive embedded platform, making it well suited for practical in-vehicle safety applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/6a02c2fdce8c8c81e964046ehttps://doi.org/10.1038/s41598-026-52381-2
Ask AI
Helpful
Bookmark
Share
View Full Paper