Safety testing serves as the fundamental pillar for the development of autonomous driving systems (ADSs), and decision-making plays a key role in ADSs. To ensure the safety of ADSs, it is paramount to generate a range of critical test scenarios to test the safety of decision-making in ADSs. While existing researches primarily focus on reproducing real-world traffic accidents in simulation environments to create test scenarios, it is essential to highlight that many of these accidents do not result in safety violations of decision-making in ADSs due to the differences between human driving and autonomous driving. More importantly, we observe that some accident-free real-world scenarios can lead to misbehaviors of ADSs. Therefore, orthogonally to existing work, it is equally important to discover safety violations of ADSs from routine traffic scenarios (i.e., accident-free scenarios) to ensure the safety of Autonomous Vehicles (AVs). We introduce CRISER , a novel methodology to achieve the above goal. It automatically generates abstract and concrete scenarios from real-traffic videos where human-driving worked safely. Based on them, CRISER discovers safety violations of the ADS's decision-making in semantic equivalent scenarios (i.e., the test scenarios with the same semantics as the original accident-free traffic videos). Specifically, CRISER enhances the ability of Large Multimodal Models (LMMs) to accurately extract scenario semantics from accident-free traffic videos and generate test scenarios by multi-modal few-shot Chain of Thought (CoT). Based on them, CRISER explores the behavior differences between the ego vehicle (i.e., the vehicle connected to the ADS under test) and human-driving in semantic equivalent scenarios. During the exploration search, CRISER keeps the semantic consistency of test scenarios with accident-free traffic videos, and explores the universality of discovered safety violations of the ADS. We implement and evaluate CRISER on the industrial-grade Level-4 ADS, Apollo. The experimental results demonstrate that CRISER can accurately extract scenario semantics and generate test scenarios from traffic videos, and effectively discover distinct types of safety violations of Apollo’s decision-making in accident-free traffic scenarios.
Tian et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: