Key points are not available for this paper at this time.
Vehicle detection in complex traffic scenes remains challenging due to frequent occlusions, lighting variations, and extreme weather. We present Hybrid-YOLO, a real-time detection framework that unifies Mamba-based state space modeling, Transformer-driven global attention, and multi-scale feature fusion to achieve high accuracy at low computational cost. At its core, Hybrid-YOLO introduces a Dynamic Residual Stem (DR Stem) for adaptive feature calibration, a Hexa-Scan Selective Block (HSSBlock) for six-directional structural perception, and a Selective State Space Model (SSM) for efficient long-range dependency modeling. A Cross-Stage Scales Feature Extraction (CSSFE) module enriches spatial semantics for small-object detection, while a Sparse-Queries Cascade Self-Attention (SCS) module focuses computation on informative regions, enhancing robustness to clutter and background noise. Extensive experiments on KITTI, BDD100K, and IITM-HeTra show that Hybrid-YOLO achieves 90.11 mAP@0.5 at 66.3 FPS, surpassing state-of-the-art methods in both accuracy and efficiency, and offering a promising solution for real-world intelligent transportation systems.
Wang et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: