PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 2026Biomedical Physics & Engineering Express0 citationsOpen Access

Multimodal Skin Disease Classification Using Vision Transformers, Medical Captioning, and Metadata Fusion: An Analysis on the ISIC 2024 Dataset

View Full Paper
ASAadit ShresthaAPAditi Palit

Key Points

  • The central aim is to improve skin lesion classification through multimodal embeddings that integrate image and clinical metadata.
  • Evaluated MedCLIP-based multimodal embeddings for binary skin lesion classification.
  • Used early and attention-based fusion of image-text representations with patient metadata.
  • Conducted experiments with multilayer perceptron and ensemble classifiers on the ISIC 2024 dataset.
  • Achieved an accuracy of 96% and an AUROC of 0.987 for skin lesion classification.
  • Outperformed unimodal baselines, demonstrating the added value of combining language and metadata.

Abstract

Skin cancer and dermatological diseases are among the most prevalent global health conditions, where early and accurate diagnosis is critical for improving patient outcomes. Although deep learning models have achieved strong performance in dermoscopic image classification, many existing approaches primarily rely on visual features and make limited use of complementary clinical metadata and language-based context routinely considered by dermatologists. Recent vision-language models (VLMs), including medical-domain adaptations such as MedCLIP, have begun to show promise in dermatology; however, their integration with structured clinical metadata and the impact of different multimodal fusion strategies have not been systematically analyzed. In this work, we address the binary skin lesion classification problem by conducting a structured evaluation of MedCLIP-based multimodal embeddings combined with classical machine learning and neural classifiers. Image-text representations are extracted using MedCLIP and fused with patient metadata through early and attention-based fusion mechanisms, followed by multilayer perceptron (MLP) and ensemble classifiers. Experiments are performed on a curated subset of the ISIC 2024 dataset comprising 1,600 training and 400 test dermoscopic images with associated metadata. The proposed multimodal approach achieves an accuracy of 96% (95.7% exact) with AUROC = 0.987, outperforming unimodal baselines and demonstrating the complementary value of language and metadata for skin lesion diagnosis. This study provides a comprehensive analysis of MedCLIP-based multimodal learning in dermatology and highlights the importance of fusion design in vision-language-metadata systems for computer-aided diagnosis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shrestha et al. (2026) studied this question.

synapsesocial.com/papers/69b25adb96eeacc4fcec8f03https://doi.org/10.1088/2057-1976/ae4eeb
Ask AI
Helpful
Bookmark
Share
View Full Paper