Key points are not available for this paper at this time.
Timely and accurate classification of medical eye images is important for supporting early screening and computer-aided decision support. However, automated ocular image classification remains challenging. Traditional convolutional neural networks (CNNs) are effective in capturing local visual features but may have limited ability to model broader spatial relationships. On the other hand, Vision Transformer-based models can capture global dependencies but may require complementary local feature extraction to identify subtle image-level patterns. To overcome these challenges, we propose MedViTNet, a hybrid deep learning framework that integrates two specialized modules: EyeConvNet for localized feature extraction and VisionNet for global contextual representation. MedViTNet was evaluated on three publicly available datasets, Myopia, Retina, and Jaundice, achieving classification accuracies of 99.99%, 99.93%, and 95.65%, respectively. Moreover, interpretability is enhanced using Grad-CAM visualizations, which highlight image regions that contribute to MedViTNet classification decisions, thereby supporting transparent and explainable computer-aided medical eye-image analysis. The results demonstrate that MedViTNet provides a unified local–global feature learning framework for heterogeneous medical eye-image classification across multiple binary classification tasks. Future work will enhance the dataset to represent a wider variety of medical cases and adapt the framework for multiclass classification tasks.
Danish et al. (Thu,) studied this question.