PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 22, 2026Intelligence-Based Medicine0 citationsOpen Access

STViTDA-Net: An explainable transformer-based framework with STGAN-ViT-MAE and deformable attention for multi-class skin cancer classification

View Full Paper
RJRavula jyothsnaKPK. PrasannaUMU. Moulali

Key Points

Key points are not available for this paper at this time.

Abstract

Skin cancer continues to pose a major global health challenge, and its early identification is essential for improving patient outcomes. Traditional diagnostic practices rely heavily on clinician expertise and manual interpretation of dermoscopic images, making the process subjective, inconsistent, and time-consuming. To address these limitations, this work introduces STViTDA-Net, an explainable transformer-based framework designed for fast, objective, and scalable multi-class skin cancer classification. The model integrates three key components: STGAN for class-balanced dermoscopic image augmentation, ViT-MAE for robust hierarchical feature learning through masked patch reconstruction, and a Deformable Attention Transformer Encoder that adaptively focuses on irregular lesion boundaries and subtle spatial variations. Preprocessing with Error Level Analysis (ELA) enhances fine-grained diagnostic cues, while Grad-CAM provides interpretable heatmaps that highlight the regions influencing the model’s predictions. Unlike manual dermoscopic evaluation, STViTDA-Net performs end-to-end inference within milliseconds and delivers consistent, expert-independent predictions supported by visual explanations. When evaluated on the ISIC2019 dataset comprising nine lesion categories, the model achieves 99.35% accuracy, 99.0% precision, 99.5% recall, 99.2% F1-score, and 99.2% AUC-ROC, surpassing existing CNN and transformer baselines. By unifying class-balanced augmentation, adaptive feature encoding, deformable attention, and explainable outputs, STViTDA-Net establishes a powerful and efficient solution for automated dermatological diagnosis. • This study proposes STViTDA-Net, a hybrid transformer-based deep learning framework combining STGAN, ViT-MAE, and Deformable Attention for accurate and explainable multi-class skin cancer classification. • STGAN is employed for class-balanced image augmentation, effectively mitigating class imbalance in the ISIC2019 dataset through realistic minority-class lesion synthesis. • ViT-MAE enables self-supervised hierarchical feature encoding by reconstructing masked image patches, capturing both local and global lesion characteristics. • Deformable Attention within the Transformer Encoder dynamically adapts to irregular lesion geometries, enhancing spatial modeling and improving classification accuracy. • Error Level Analysis (ELA) preprocessing and Grad-CAM post-hoc explanations improve interpretability by highlighting diagnostically significant regions, supporting clinical validation. • STViTDA-Net achieves state-of-the-art performance on ISIC2019 with 99.35% accuracy, 99.0% precision, 99.5% recall, 99.2% F1-score, and 99.2% AUC-ROC, surpassing CNN, ViT, and hybrid baselines. • With its superior accuracy, robustness to class imbalance, and visual explainability, STViTDA-Net offers a scalable and clinically applicable solution for dermatological diagnostics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

jyothsna et al. (2026) studied this question.

synapsesocial.com/papers/6a0d9f40d266b659c409b91ehttps://doi.org/10.1016/j.ibmed.2026.100350
Ask AI
Helpful
Bookmark
Share
View Full Paper