PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025Frontiers in Medicine0 citationsOpen Access

A dual attention and multi-scale fusion network for diabetic retinopathy image analysis

View Full Paper
MZMenglin ZhangQLQi LiuJZJialei Zhan

Key Points

  • DWAM-MSFINET achieved 82.59% Top-1 accuracy, outperforming previous models in diabetic retinopathy image analysis.
  • The architecture includes a multi-scale feature integration module that enhances the identification of pathological patterns.
  • With a processing rate of 62.5 images per second, this method supports scalable real-time medical diagnosis.
  • Robustness against domain shift was demonstrated on the Messidor dataset, ensuring better generalization across image variations.

Abstract

Robust classification of medical images is crucial for reliable automated diagnosis, yet remains challenging due to heterogeneous lesion appearances and imaging inconsistencies. We introduce DWAM-MSFINET (Dual Window Adaptation and Multi-Scale Feature Integration Network), a novel deep neural architecture designed to address these complexities through a dual-pathway integration of attention and resolution-aware representation learning. Specifically, the Multi-Scale Feature Integration (MSFI) module hierarchically aggregates semantic cues across spatial resolutions, enhancing the network’s capacity to identify both fine-grained and coarse pathological patterns. Complementarily, the Dual Weighted Attention Mechanism (DWAM) adaptively modulates feature responses in both spatial and channel dimensions, enabling selective focus on clinically salient structures. This unified framework synergizes localized sensitivity with global semantic coherence, effectively mitigating intra-class variability and improving diagnostic generalization. DWAM-MSFINET achieved 78.6% Top-1 accuracy on the standalone Messidor dataset, demonstrating robustness against domain shift. DWAM-MSFINET surpasses state-of-the-art CNN and Transformer-based models, achieving a Top-1 accuracy of 82.59%, outperforming ResNet50 (81.68%) and Swin Transformer (80.26%), while inference latency is 16.0 ms per image (not seconds) when processing batches of 16 images on NVIDIA RTX 3090, equivalent to 62.5 images per second. These results validate the efficacy of our approach for scalable, real-time medical image analysis in clinical workflows. We have released our code and datasets at: https://github.com/eleen7/data .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68c199e29b7b07f3a061b4cfhttps://doi.org/10.3389/fmed.2025.1614046
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Optical Coherence Tomography-Angiography of a large retinal microaneurysm2020 · 3 citations
  2. 2STMSF: Swin Transformer with Multi-Scale Fusion for Remote Sensing Scene Classification2025 · 32 citations
  3. 3Prediction of Incident Diabetic Retinopathy in Adults With Type 1 Diabetes Using Machine Learning Approach: An Exploratory Study2024 · 29 citations
  4. 4Current Treatments for Diabetic Macular Edema2023 · 89 citations
  5. 5Short-term-outcomes of idiopathic epiretinal membranes treated with pars-plana-vitrectomy – examination of visual function and OCT-morphology2023 · 4 citations