Extracting critical scenarios from the vast and complex driving environments is a crucial step in the iterative upgrade process of autonomous driving systems. Existing image–text‐based retrieval methods provide a good understanding of the environment but overlook the critical information regarding the impact of the environment on the self‐vehicle. To address these issues, we propose a contrastive image–text–motion retrieval (CITMR) cross‐modal learning framework that uses descriptive text as input to retrieve critical scenarios. The framework employs image, text, and motion encoders to extract features from different modalities and uses contrastive loss to enable cross‐modal information comparison and interaction. Finally, the vehicle’s motion and its textual description are supplemented in the collected autonomous driving dataset, and a scenario dataset containing image–text–motion pairs is constructed for model validation. Experimental results show that CITMR achieves retrieval performance for text‐critical scenarios of 0.8684 and 0.8557 on two test datasets, outperforming the baseline methods.
Peng et al. (Thu,) studied this question.