Review examines visible-infrared image fusion techniques and datasets, improving object detection and facial-expression recognition applications.
Visible and infrared (IR) image fusion has become an important strategy for improving computer vision performance under low illumination, occlusion, and some poor-visibility conditions. By integrating complementary textural information from visible images with thermal or IR cues, VIR fusion can enhance object localization, detection robustness, and facial-expression recognition (FER). This review examines VIR fusion techniques and datasets for computer vision applications, with object detection (OD) considered as a relatively mature scene-level task and FER considered as an emerging human-centered application. It summarizes major multimodal datasets, compares early-fusion approaches, including sensor- and feature-level fusion, with late-fusion approaches, including score- and decision-level fusion, and discusses representative machine learning and deep learning methods. The review also evaluates commonly used performance metrics and identifies current limitations, including dataset imbalance, sensor misalignment, limited demographic diversity in facial-expression datasets, computational complexity, and weak real-time generalization. Finally, key application areas, including surveillance, healthcare, remote sensing, autonomous systems, and human–computer interaction, are discussed. This review highlights the need for better-aligned multimodal datasets, standardized evaluation protocols, lightweight fusion architectures, and robust models capable of operating in dynamic real-world environments.
No takes yet. Share an insight, caveat, or question.
Naseem et al. (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: