Underwater sonar object detection is challenging because targets are often small, boundaries are blurred, background clutter is strong, and labeled sonar data are limited. To address these issues, we propose T2C-DETR, a detector built on RT-DETR with three task-oriented improvements: (i) a Transformer–Convolution dual-channel backbone (TCDCNet) for complementary global-context and local-detail modeling, (ii) a Noise Filtering Module (NFM) inserted before neck fusion to suppress noise-dominated activations, and (iii) a stage-wise transfer-learning strategy tailored to small sonar datasets. We evaluate the method under three pre-training sources (COCO 2017, DOTA, and an infrared dataset) and then fine-tune on a self-built sonar dataset. Experimental results show that T2C-DETR achieves AP50 of 97.8%, 98.2%, and 98.5% at 72–73 FPS, consistently outperforming the RT-DETR baseline, YOLOv5-Imp, and MLFFNet in the accuracy–speed trade-off. These results indicate that combining global–local representation learning with targeted noise suppression is effective for practical real-time sonar detection.
Wu et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: