Los puntos clave no están disponibles para este artículo en este momento.
Pedestrians’ intention to cross the street exercises a substantial influence on the decision-making process of autonomous vehicles in urban traffic environments. However, accurately predicting pedestrian crossing intention is non-trivial due to the interweaving of pedestrian personalities and traffic scene elements. Despite the significant achievement of previous studies, challenges remain in effectively extracting and integrating diverse features from different modalities of observation data. In response, this paper proposes a novel model leveraging pedestrian bounding boxes, poses, and ego-vehicle speed to predict crossing intention. We introduce mixture expert feature embedding (MEFE) to project raw data into high-dimensional space based on different types of inputs. A multi-branch spatial and temporal graph convolutional network (MB-STGCN) is applied to capture multi-scale spatial and temporal features of pedestrian pose skeleton joints. A multi-token temporal aggregation (MTTA) method is devised to preserve abundant temporal information of observation data. Additionally, a progressive multimodal feature fusion (PMFF) method based on symmetric channel split attention (SCSA) is employed to enhance the interaction between different modalities when integrating different features. Extensive comparison experiments and ablation studies on public benchmark datasets substantiate the effectiveness of our approach, showing significant improvements over existing models.
Chen et al. (Fri,) studied this question.