The introduction of pioneering technologies such as programmable gradient information and Generalized Efficient Layer Aggregation Network in YOLOv9 has significantly improved its efficiency, accuracy, and adaptability compared to YOLOv8. The proposal of SPPELAN has notably enhanced the network's feature extraction capability. However, SPPELAN's use of multi-level maximum pooling layers may lead to the loss of some detail information, especially with larger pooling kernel sizes, potentially ignoring smaller targets or details. To address this, we propose the Spatial Multi- Fusion Net, which involves segmenting channels and then blending image features extracted from different channels using maximum pooling blocks with different kernel sizes and convolutional blocks of varying depths. This allows the model to capture features at different abstraction levels, thereby achieving the goal of collecting features of objects of different sizes. Integrating the Spatial Multi-Fusion Net into YOLOv9 further improves its performance on the COCO dataset's object detection task, with all metrics showing enhancement, despite adding fewer parameters.
No takes yet. Share an insight, caveat, or question.
Zizhuang Liu (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: