PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 21, 2026Bioengineering0 citationsOpen Access

Multimodal Deep Learning with Attention-Based Fusion for Skin Cancer Diagnosis

View Full Paper
WAWiem AbdelbakiHAHend AlshayaINInzamam Mashood Nasir

Key Points

  • This research aims to develop a multimodal deep learning framework to improve the diagnosis of skin cancer by integrating clinical information with dermoscopic image features.
  • A multimodal deep learning architecture was proposed, which combines attention-based feature encoding with structured fusion of images and clinical data.
  • Evaluations were conducted on benchmark datasets including ISIC 2019, ISIC 2020, and HAM10000 with a unified experimental design.
  • Ablation analysis was performed to test the effectiveness of attention and fusion components.
  • The framework achieved accuracies of 90.5%, 88.7%, and 91.8% and AUCs of 95.8%, 94.6%, and 96.3% on the respective datasets.
  • Increased AUC by 6.5% and F1 score by 8.0% compared to baseline models (ResNet50 and EfficientNet-B4).
  • AUC improvement was statistically significant across all metrics (p = 0.01).

Abstract

The diagnosis of skin cancer remains a growing challenge because of its high variability as a result of the varying imaging conditions in clinical settings. This paper proposes a multimodal deep learning framework to address these challenges by combining the auxiliary clinical information with dermoscopic image features. This proposed architecture uses an attention-based feature encoder with a structured multimodal fusion approach to utilize the integrated feature representation across all channels. Evaluations of the proposed architecture were conducted across a range of benchmark datasets, including ISIC 2019, ISIC 2020, and HAM10000, using a unified experimental approach. This proposed model achieved accuracies of 90.5%, 88.7%, and 91.8% and AUCs of 95.8%, 94.6%, and 96.3%, respectively, on the selected datasets. For the baseline models, ResNet50 and EfficientNet-B4, our approach increased the AUC by 6.5% and the F1 score by 8.0%. Furthermore, across various datasets, the model achieved an AUC of 90.9%, proposing strong generalization. From the ablation analysis results, the attention and multimodal fusion mechanisms showed a 4.1% decrease in AUC when key components were removed, confirming their effectiveness. With 34.7 million parameters and an average of 19.3 Ms., the model has adequate intensity to deploy in a real clinical setting without affecting its performance. Additionally, the improvements to the model were statistically significant across all evaluation metrics (p = 0.01). The proposed multimodal framework demonstrates strong performance and robustness across multiple benchmark datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abdelbaki et al. (2026) studied this question.

synapsesocial.com/papers/6a0ea15cbe05d6e3efb5fe14https://doi.org/10.3390/bioengineering13050564
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparative analysis of multimodal architectures for effective skin lesion detection using clinical and image data2025 · 5 citations
  2. 2A Hybrid Fusion Approach for Skin Cancer Detection using Deep Learning on Clinical Images and Machine Learning on Patient Metadata2025
  3. 3Ensemble-Based Multimodal Deep Learning for Precise Skin Cancer Diagnosis: Integrating Clinical Imagery with Patient Metadata2026
  4. 4Multi‐Scale Attention Fusion With Depthwise Separable Convolutions for Efficient Skin Cancer Detection2025
  5. 5<p>Multimodal Skin Disease Classification via Gated Cross-Modal Fusion of Dermoscopic Images and Clinical Metadata</p>2026