Multimodal data fusion has emerged as a promising approach for improving the accuracy and robustness of diabetic retinopathy (DR) detection by integrating complementary information from heterogeneous data sources. This review systematically examines recent deep learning-based multimodal approaches for DR detection, focusing on model architectures, modality combinations, fusion strategies, and reported performance outcomes. The analysis reveals that while convolutional neural networks (CNNs) remain dominant, hybrid architectures are increasingly adopted to handle heterogeneous inputs. Combinations such as fundus with OCT or electronic health records (EHR) are frequently associated with improved performance, although results remain highly dependent on dataset characteristics and experimental design. Across studies, early and joint fusion strategies are commonly applied, but no single approach consistently outperforms others, highlighting the context-dependent nature of fusion design. Overall, this review identifies key trends, methodological gaps, and emerging directions, providing a clearer foundation for developing robust and clinically applicable multimodal DR detection systems.
Wardhani et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: