This paper introduces autonomous driving image perception technology, including deep learning models (such as CNN and RNN) and their applications, analyzing the limitations of traditional algorithms. It elaborates on the shortcomings of Faster R-CNN and YOLO series models, proposes various improvement techniques such as data fusion, attention mechanisms, and model compression, and introduces relevant datasets, evaluation metrics, and testing frameworks to demonstrate the advantages of the improved models.
Guanglei Xiong (2025) studied this question.