PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 26, 2026Cutaneous and Ocular Toxicology0 citations

Performance of multimodal Large Language Models in misdiagnosed dermatologic cases: a pilot study on diagnostic accuracy and human error replication

View Full Paper
OŞOrhan Şen

Key Points

  • To assess the diagnostic accuracy of multimodal large language models in clinically ambiguous dermatologic cases.
  • Cross-sectional analysis on 30 biopsy-confirmed diagnostic dilemmas.
  • Models evaluated using text-only and multimodal queries.
  • Primary outcomes included Top-1 accuracy and replication rates of human error.
  • Gemini 3 achieved the highest multimodal Top-1 accuracy at 60.0% (18/30).
  • In the inflammatory subgroup, Gemini 3's accuracy improved from 45.5% (5/11) in text-only to 72.7% (8/11) in multimodal mode; not significant, p = 0.248.
  • All models demonstrated limited accuracy for malignant lesions using macro-images.

Abstract

BACKGROUND: Multimodal Large Language Models (LLMs) are increasingly positioned as diagnostic assistants in dermatology. However, current research often relies on clear-cut cases, leaving their performance in clinically ambiguous, gray zone scenarios insufficiently explored. Specifically, whether integrating visual data helps LLMs correct initial human misdiagnoses or reinforces cognitive biases remains unknown. OBJECTIVES: To evaluate the diagnostic accuracy of three recent multimodal large language models all queried through their default web interfaces on 5 February 2026, using standardized single-turn prompts in biopsy-confirmed dermatologic cases initially misdiagnosed by clinicians, and to assess the impact of visual integration on human error replication rates. METHODS: A cross-sectional analysis was conducted on 30 diagnostic dilemmas confirmed by histopathology. Models were queried using a two-stage protocol: (1) Text-Only and (2) Multimodal. Primary outcomes were Top-1 accuracy, visual gain, and the rate of replicating the clinician's initial error. RESULTS: Gemini 3 achieved the highest multimodal Top-1 accuracy (60.0% (18/30), followed by ChatGPT-5.2 at 56.7% (17/30) and Claude 4.5 Sonnet at 33.3% (10/30). In the inflammatory subgroup, Gemini 3 accuracy increased from 45.5% (5/11) in text-only to 72.7% (8/11) in multimodal mode; this difference was not statistically significant (McNemar's test, p = 0.248). All models showed limited accuracy for malignant lesions using macro-images. CONCLUSIONS: While Gemini 3 shows promise as a de-biasing tool in complex inflammatory dermatoses, multimodal LLMs currently lack the granular precision required for malignancy detection without dermoscopic data. These findings underscore the need for cautious integration of AI in high-stakes diagnostic scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Orhan Şen (2026) studied this question.

synapsesocial.com/papers/6a3e1907030ad1a9b3091e01https://doi.org/10.1080/15569527.2026.2692370
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Multimodal large language models for oral lesion diagnosis: a systematic review of diagnostic performance and clinical utility2026 · 12 citations
  2. 2Exploratory Evaluation of Diagnostic Accuracy and Temporal Reproducibility of Multimodal Large Language Models in the Image-Based Assessment of Oral Mucosal Lesions2026
  3. 3Comparative Analysis of Large Language Models in Dermatological Diagnosis: An Evaluation of Diagnostic Accuracy2025 · 3 citations
  4. 4Multimodal Large Language Models in Ophthalmology: Diagnostic Accuracy and the Risk of Metadata-Induced Confirmation Bias2026
  5. 5Comparison of Multimodal Large Language Models and Physicians for Medical Diagnosis Using NEJM Image Challenge Cases: Cross-sectional Study2025