Visible and infrared image fusion is a feasible solution to improve object detection in poor lighting scenes. Existing methods first fuse visible and infrared images to obtain fused images, and then fed into object detection models trained on datasets with only visible images. The two-stage approach leads to a large parameter quantity and computation complexity. We constructed a multispectral dataset and investigated end-to-end models through multispectral image fusion to improve object detection performance while maintaining fewer parameters and complexity. Distinct from existing datasets, our multispectral dataset focuses on road traffic scenarios under extreme climatic and lighting conditions, with higher imaging quality and resolution. Pixel alignment of visible and infrared images was accomplished via coaxial imaging and homography transformation with small rotation and translation, achieving registration error within three pixels. Two types of end-to-end object detection models for visible and infrared image fusion, named 4-channel and Y-shaped architectures, are proposed. Experimental results show that both the 4-channel and Y-shaped models improve object detection accuracy in all scenes. Feature fusion timing and attention mechanisms were studied to further improve the detection performance. The experiments show that the comprehensive performance is optimal when the visible and infrared image features are fused in the middle stage, and the attention mechanisms allow to extract relevant and complementary features of visible and infrared images. Our models achieve better performance with fewer parameters and lower computation complexity compared with the SOTA.
Zhao et al. (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: