Mainstream real-time object detectors, like the YOLO series, balance speed and accuracy but are bottlenecked by Non-Maximum Suppression (NMS) for post-processing. While end-to-end Transformer-based detectors show potential by eliminating NMS, their high computational cost impedes real-time application. Concurrently, alternatives like DECO, using a pure convolutional framework, are hampered by local receptive fields, limiting global context modeling and limiting the accuracy on lightweight models. To address this, we introduce HA-DETR, a Hybrid Architecture DETR fusing the local feature extraction of convolutions with the global context modeling of Transformers, particularly effective in resource-constrained scenarios. HA-DETR uses an efficient multi-scale encoder and an effective hybrid decoder that integrates convolutional query refinement with cross-attention to accelerate predictions. This hybrid design, however, exacerbates a training dynamic mismatch. To mitigate this, we propose the Decoupled Gamma Loss (DGL), which introduces independent modulating factors, ₎ₒ and ₍₄₆, to alleviate sample imbalance from disparate convergence rates of the hybrid components. Experiments validate our method. On the COCO dataset, our lightweight HA-DETR-R18 achieves 48. 4 AP at 68 FPS on a V100 GPU, surpassing RT-DETR-R18 by 1. 9 AP with a 13% speedup and outperforming DECO-R18 by 7. 9 AP. Our work provides a competitive architectural alternative for efficient and accurate real-time object detection.
Tang et al. (Thu,) studied this question.