Abstract Achieving accurate detection of small and occluded objects in complex natural scenes remains challenging for real-time systems. This paper presents CSS-YOLO, a lightweight detector built upon YOLOv11 that enhances multi-scale semantic consistency, channel discriminability, and local spatial modeling. CSS-YOLO integrates a cross-channel feature fusion module, a multi-branch SENetV2 channel attention, and a spatially coupled convolution block. On a self-constructed Rosa davurica Pall dataset, CSS-YOLO reaches a mean average precision (mAP) at Intersection over Union (IoU) 0.5 (mAP@0.5) of 93.5% and mAP@0.5:0.95 of 64.0%, while reducing parameters by 33.1%, Giga Floating-point Operations Per Second (GFLOPs) by 17.5%, and model size by 30.8% compared with YOLOv11 . Ablation studies quantify the contribution of each module; comparisons on pomegranate and strawberry datasets demonstrate strong cross-scenario generalization. The results indicate that CSS-YOLO offers a practical balance of accuracy and efficiency suitable for edge deployment in smart agriculture and similar real-time applications.
Han et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: