This is the authors' abstract. We don't add key points for this paper.
The Transformer model [1] [6] [7], with its selfattention mechanism and global context modeling capability, has introduced a new paradigm for object detection. Unlike traditional methods that rely on anchor-box design and local convolution, the Transformer-based approach, with its end-to-end architecture and long-range dependency modeling, has propelled object detection from experience-driven local optimization to data-driven global perception. This paper systematically reviews the research progress of Transformer-based object detection, focusing on the following aspects: technical framework innovation, multi-scenario application adaptation, challenge and trend balancing. By integrating algorithm evolution, scenario verification, and systematized evaluation criteria (such as the COCO dataset and mAP metrics), this review provides a theoretical framework and cross-domain reference for the deepened application of Transformers in object detection.
No takes yet. Share an insight, caveat, or question.
Liu et al. (2025) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: