PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 5, 2026JAMA Dermatology1 citations

Limits of Artificial Intelligence Models for Skin Cancer Diagnosis in Realistic Settings

View Full Paper
JAJulien AnriotSYSiyuan YanCCClio Coste

Key Points

  • This study aims to evaluate and compare the diagnostic accuracy of AI algorithms against human evaluators with varying levels of expertise in skin cancer diagnosis.
  • Multi-institutional diagnostic study comparing AI models and physician readers for skin lesion diagnosis.
  • Dataset of 1117 dermatological images used, evaluated by 652 physicians across different experience levels.
  • Primary outcome was multiclass diagnostic accuracy; secondary outcomes included sensitivity and specificity.
  • All human readers outperformed the CNN with an accuracy of 65.9% vs 56.7%, P<0.001.
  • Unimodal AI accuracy (72.2%) surpasses human readers with less than 3 years of experience (68.2%), P<0.001.
  • Experts (≥10 years) achieved the highest accuracy (74.2%), outperforming all AI models.

Abstract

Importance: Artificial intelligence (AI) systems for skin cancer detection perform well in controlled settings but frequently underperform in everyday clinical practice, raising critical questions about their readiness for deployment. Objective: To compare the diagnostic accuracy of AI algorithms vs human evaluators across varying expertise levels for skin lesion diagnosis, including rare and atypical cases, in a realistic clinical context. Design, Setting, and Participants: This multi-institutional diagnostic study compared diagnostic performance among AI models and physician readers with varying dermatological expertise, ranging from less than 1 year to more than 10 years of experience. A dataset of dermatological images representing everyday clinical scenarios was used and contained 1117 cases, including clinical and dermoscopic images with associated metadata. Study inclusion spanned from March 16, 2023, to August 1, 2025. Exposures: Three AI algorithms: a first-generation convolutional neural network (CNN) and 2 foundation models (PanDerm unimodal and multimodal). Human readers evaluated 100 stratified, random cases drawn from the same dataset. Main Outcomes and Measures: The primary outcome was reader-level multiclass diagnostic accuracy for skin lesion classification. Secondary outcomes were binary benign vs malignant sensitivity, specificity, and balanced accuracy. Performance was compared between AI algorithms and human readers stratified by experience level. Results: A total of 652 physicians (median IQR age, 33 29-37 years; 559 85.7% female) contributed to 1092 test iterations. All human readers outperformed the CNN (mean SD accuracy, 65.9% 10.5% vs 56.7% 3.9%; difference, 9.2 percentage points pp; 95% CI, -9.8 to 8.5 pp; P < .001). Unimodal accuracy exceeded readers with less than 3 years of experience (mean SD accuracy, 72.2% 3.5% vs 68.2% 7.6%; difference, 4.0 pp; 95% CI, 3.2-4.9 pp; P < .001). With a mean (SD) accuracy of 74.2% (5.7%), experts with more than 10 years of experience achieved the highest multiclass diagnostic accuracy, outperforming all AI models on this primary end point, which included 56.7% (3.9%) for CNN, 72.2% (3.5%) for the unimodal model, and 66.3% (3.8%) for the multimodal model. Conclusions and Relevance: In this diagnostic study, a modern foundation model surpassed readers with less than 3 years of experience on accuracy of skin lesion diagnosis and matched those with 3 to 10 years of experience but remained inferior to experts with more than 10 years of experience, highlighting both the promise and current limitations of AI in dermatologic diagnosis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Anriot et al. (2026) studied this question.

synapsesocial.com/papers/6a22688f763171746d547165https://doi.org/10.1001/jamadermatol.2026.1492
Ask AI
Helpful
Bookmark
Share
View Full Paper