PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 12, 2025Cureus3 citationsOpen Access

Comparative Analysis of Large Language Models in Dermatological Diagnosis: An Evaluation of Diagnostic Accuracy

View Full Paper
NTNiharika TekchandaniAMArchana MukherjeeNPNandakumar Poonthottam

Key Points

  • ChatGPT-4o and Claude 3.7 correctly identified the top diagnosis in 66.7% of cases, outperforming Gemini 2.0 Flash.
  • Assessment of three AI models showed ChatGPT-4o and Claude 3.7 achieved 86.7% total coverage in differential matches versus 60.0% for Gemini 2.0.
  • Retrospective analysis utilized 15 case reports to evaluate diagnostic outputs of AI models based on clinical presentations.
  • Findings underscore the need for continued refinement and validation of LLMs as adjunctive tools in dermatology.

Abstract

Background: The diagnostic process in dermatology often hinges on visual recognition and clinical pattern matching, making it an attractive field for the application of artificial intelligence (AI). Large language models (LLMs) like ChatGPT-4o, Claude 3.7 Sonnet, and Gemini 2.0 Flash offer new possibilities for augmenting diagnostic reasoning, particularly in rare or diagnostically challenging cases. This study evaluates and compares the diagnostic capabilities of these LLMs based solely on clinical presentations extracted from rare dermatological case reports. Methodology: Fifteen published case reports of rare dermatological conditions were retrospectively selected. Key clinical features, excluding laboratory or histopathological findings, were input into each of the three LLMs using standardized prompts. Each model produced a most probable diagnosis and a list of differential diagnoses. The outputs were evaluated for top-match accuracy and whether the correct diagnosis was included in the differential list. Performance was analyzed descriptively, with visual aids (heatmaps, bar charts) illustrating comparative outcomes. Results: ChatGPT-4o and Claude 3.7 Sonnet each correctly identified the top diagnosis in 10 (66.7%) out of 15 cases, compared to 8 (53.3%) out of 15 for Gemini 2.0 Flash. When differential-only matches were included, both ChatGPT-4o and Claude 3.7 achieved a total coverage of 86.7%, while Gemini 2.0 reached 60.0%. Notably, all models failed to identify certain diagnoses, including blastic plasmacytoid dendritic cell neoplasm and amelanotic melanoma, underscoring the potential risks associated with plausible but incorrect outputs. Conclusions: This study demonstrates that ChatGPT-4o and Claude 3.7 Sonnet show promising diagnostic potential in rare dermatologic cases, outperforming Gemini 2.0 Flash in both accuracy and diagnostic breadth. While LLMs may assist in clinical reasoning, particularly in settings with limited dermatology expertise, they should be used as adjunctive tools, not substitutes, for clinician judgment. Further refinement, validation, and integration into clinical workflows are warranted.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tekchandani et al. (2025) studied this question.

synapsesocial.com/papers/68d44a1d31b076d99fa52eb4https://doi.org/10.7759/cureus.92089
Ask AI
Helpful
Bookmark
Share
View Full Paper