PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026Frontiers in Artificial Intelligence0 citationsOpen Access

An automated framework to classify skin lesions using Multi-Head Self Attention Layer-based Vision Transformers

SFSahil FaizalCRCharu Anant RajputMPManas Ranjan Prusty

Key Points

  • The study aims to develop an automated system for classifying skin lesions into nine distinct categories to aid in early detection of malignancies.
  • Utilized contrast stretching for image enhancement and region of interest (ROI) segmentation.
  • Implemented a Vision Transformer (ViT) for feature extraction from skin lesion images.
  • Employed a light-weight multi-layer perceptron (MLP) for multinomial classification of lesions.
  • Achieved training accuracy of 98% and testing accuracy of approximately 93.22%.
  • Successfully classified skin lesions into nine categories, including squamous cell carcinoma and melanoma.
  • Demonstrated scalability for extension to additional diagnostic classes in future research.

Abstract

Skin lesions are one of the most prevalent form of diseases existing among us. Early detection and classification of potentially malignant skin lesions can give us a lead in the fight against skin cancer. There are many lesion classification divisions on medical grounds; however, an automated system that detects and classifies a majority of these classes is not prevalent. In view of this scenario, our proposed study aims to classify the input skin lesion images into nine classes, namely, squamous cell carcinoma (SCC), Basal cell carcinoma (BCC), melanocytic nevi (NV), actinic keratoses and intraepithelial carcinoma (AKIEC), melanoma (MEL), seborrheic keratosis (SEK), dermatofibroma (DF), benign keratosis like lesions (BKL), and vascular lesions (VASC). The proposed methodology uses contrast stretching as an image enhancement technique to facilitate efficient Region of Interest (ROI) segmentation. The novelty of the proposed study lies in the first-hand implementation of Vision Transformer (ViT) for feature extraction in the domain of skin lesion detection. Finally, a light-weight multi-layer perceptron (MLP) composed of fully connected layers is used for multinomial classification. Combining the aforementioned techniques, the proposed method achieves training accuracy of 98% and testing accuracy of about 93.22%. The impressive performance across nine distinct categories represents a significant milestone. This success demonstrates the model’s scalability, suggesting it can be effectively extended to a broader array of diagnostic classes in future research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Faizal et al. (2026) studied this question.

synapsesocial.com/papers/69fece83b9154b0b82875e91https://doi.org/10.3389/frai.2026.1781796
Ask AI
Helpful
Bookmark
Share
View Full Paper