PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 26, 2026Bioengineering4 citationsOpen Access

Dual-SwinOrd: A Dual-Head Swin Transformer with Semantic Prior Injection for Ordinal Diabetic Retinopathy Grading

View Full Paper
WYWenjuan YuXSXiaonan SiJZJingxiang Zhong

Key Points

  • This research aims to develop an effective framework for grading diabetic retinopathy by integrating deep learning and semantic guidance.
  • Developed Dual-SwinOrd framework using a Swin Transformer backbone for feature extraction.
  • Incorporated Progressive Lesion-aware Kernel Attention for diverse lesion scale handling.
  • Utilized Semantic Prior Modulation guided by PubMedCLIP to align visual features with medical knowledge.
  • Implemented a Dual-Head learning strategy for simultaneous classification and ordinal regression.
  • Achieved 87.98% accuracy and 0.9370 QWK on the APTOS 2019 dataset.
  • Achieved 86.54% accuracy and 0.9040 QWK on the DDR dataset.
  • Demonstrated improvements in capturing long-range relationships and semantic reasoning compared to standard models.

Abstract

Diabetic retinopathy (DR) is the largest cause of permanent vision loss in the working-age population, making automated grading critical for timely therapeutic intervention. While recent deep learning algorithms have improved feature discrimination, modern state-of-the-art systems have two fundamental drawbacks. First, most models rely on standard Convolutional Neural Networks, which struggle to capture long-range relationships and lack semantic reasoning, resulting in visual findings that do not correlate with clinical knowledge. Second, present approaches often consider grading as a nominal classification or a pure ordinal regression task, failing to strike a compromise between high classification accuracy and severity-consistent predictions (Quadratic Weighted Kappa). To address these challenges, we propose Dual-SwinOrd, a novel framework that integrates a hierarchical Vision Transformer with a semantically guided dual-head mechanism. Specifically, we use a Swin Transformer backbone to extract hierarchical features, effectively capturing global retinal structures. To handle diverse lesion scales, we incorporate a Progressive Lesion-aware Kernel Attention (PLKA) module and a Semantic Prior Modulation (SPM) module guided by PubMedCLIP, bridging the gap between visual features and medical linguistic priors. In addition, we propose a Dual-Head learning strategy that decouples the optimization objective into two parallel streams: a Classification Head to maximize diagnostic accuracy and an Ordinal Regression Head (DPE) to enforce rank-consistency. This design effectively mitigates the trade-off between precision and ordinality. Extensive experiments on the APTOS 2019 and DDR datasets demonstrate that Dual-SwinOrd achieves state-of-the-art performance, yielding an Accuracy of 87.98% and a Quadratic Weighted Kappa (QWK) of 0.9370 on the APTOS 2019 dataset, as well as an Accuracy of 86.54% and a QWK of 0.9040 on the DDR dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yu et al. (2026) studied this question.

synapsesocial.com/papers/69c4cd30fdc3bde448919258https://doi.org/10.3390/bioengineering13040374
Ask AI
Helpful
Bookmark
Share
View Full Paper