Objective To systematically evaluate transfer learning (TL) models for multiclass ocular disease diagnosis and assess their reliability using explainable artificial intelligence (AI). Methods Eight pretrained convolutional neural network (CNN) models were evaluated on a public dataset covering cataract, diabetic retinopathy, glaucoma, and normal classes under a unified protocol. Performance was measured using accuracy, precision, recall, and F1-score. Grad-CAM, LIME, and SHAP were used for interpretability, and the Friedman test assessed performance consistency. Results Several models achieved near-perfect performance for diabetic retinopathy. DenseNet121 and XceptionNet performed best for cataract detection, while glaucoma showed consistently weaker results, indicating the need for segmentation-based approaches. Despite similar accuracy, explainability revealed substantial differences in model attention. EfficientNetB3 produced the most clinically meaningful visual explanations. Conclusions Accuracy alone is insufficient for trustworthy medical AI. Explainable AI is essential for model selection. EfficientNetB3 offers the best balance between performance and interpretability, and glaucoma diagnosis requires more advanced, segmentation-aware pipelines.
Nisa et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: