Key points are not available for this paper at this time.
Recognizing the crossing behavior of non-motorized traffic at intersections is crucial for urban transportation safety and efficiency. However, traditional methods often overlook environmental context, while deep learning approaches require extensive labeled data. To address these limitations, we propose an unsupervised learning framework that infers waiting and passing behaviors by combining motion features with spatio-temporal semantic context. Our model employs a convolutional autoencoder with an integrated clustering layer, trained through a joint optimization of reconstruction, clustering, and K-neighborhood contrastive losses. Experimental results show our framework outperforms other unsupervised methods. Notably, incorporating semantic features improves clustering accuracy from 87.3% to 91.3%, validating the effectiveness of our approach in capturing complex behavior.
Xu et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: