Multimodal deep learning classifies diabetic retinopathy severity in large datasets, indicating high performance and interpretability for clinical use.
Diabetic retinopathy remains a leading cause of vision loss worldwide, with early detection essential to prevent irreversible impairment. This study proposes an explainable multimodal deep learning framework that integrates high-resolution retinal fundus images with structured clinical metadata for diabetic retinopathy severity classification. The architecture combines convolutional features extracted using EfficientNetB0 with a parallel multi-layer perception trained on patient-specific variables including age, diabetes duration, HbA1c, blood pressure, and body mass index. On the APTOS 2019 dataset (n = 3,662), the model achieved a 5-fold cross-validation accuracy of 81.3%, a weighted F1-score of 0.80, and an AUC of 0.936 across five severity stages. External validation on Messidor-2 demonstrated performance degradation, indicating domain shift and limited cross-dataset generalization. Grad-CAM visualizations highlighted clinically relevant retinal regions, while SHAP analysis confirmed the importance of systemic risk factors such as HbA1c and diabetes duration. The results indicate strong performance, interpretability, and suitability for screening and tele-ophthalmology workflows.
No takes yet. Share an insight, caveat, or question.
Rahat et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: