PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 11, 2026The Journal of Infectious Diseases3 citations

Diagnostic Accuracy of Commercial Large Language Models for Anogenital Skin Lesion Images: A Comparative Study of Gemini, Claude and ChatGPT

View Full Paper
NSNyi Nyi SoePLPhyu Mon LattDLDavid Lee

Key Points

  • Evaluate the diagnostic accuracy of three LLMs for diagnosing anogenital dermatological conditions using clinical images.
  • Used de-identified clinical images of anogenital conditions from the STI Atlas and public sources.
  • Measured accuracy in correctly classifying images as STIs vs non-STIs and ranking diagnoses in top-1, top-3, and top-5.
  • Conducted between September and November 2025 with a sample size of 218 images.
  • Gemini achieved the highest classification accuracy at 76.2% for STI binary classification (95% CI, 70.5% - 81.9%).
  • Top-1 accuracy for Gemini was 39.0% (95% CI, 32.7% - 45.7%); top-3 was 54.6% (95% CI, 47.9% - 61.1%); top-5 was 60.6% (95% CI, 53.9% - 66.9%).
  • All LLMs struggled significantly with challenging images, showing top-5 accuracy between 29.2% and 40.0% and indicating the need for improved diagnostic reliability.

Abstract

Abstract Background Diagnosing anogenital dermatological conditions often requires specialist expertise that is unavailable in many clinical settings. Large language models (LLMs) are increasingly accessible to clinicians, but their diagnostic accuracy for anogenital dermatology has not been evaluated. We evaluated the diagnostic accuracy of three LLMs (Gemini 2.5 Pro, Claude Opus 4.1, and ChatGPT 5 Thinking). Methods This study was conducted between September and November 2025, using de-identified clinical images of anogenital conditions from the STI Atlas (stiatlas.org (https://stiatlas.org/)) and other publicly available sources. Primary outcomes were correct classification of images identified as sexually transmitted infections (STIs) vs non-STIs and the inclusion of the correct diagnosis among the LLMs’ top-ranked (top-1), top-3, or top-5 differential diagnoses. Results Among 218 images, Gemini achieved the highest accuracy for STI binary classification (76.2% 95% CI, 70.5% - 81.9%) and differential diagnosis (top-1, 39.0% 95% CI, 32.7% - 45.7%; top-3, 54.6% 95% CI, 47.9% - 61.1%; top-5, 60.6% 95% CI, 53.9% - 66.9%), followed by ChatGPT and Claude. In subgroup analysis, all LLMs showed substantially reduced accuracy for diagnostically challenging images (top-5 accuracy range, 29.2% - 40.0%). Gemini consistently outperformed Claude across most subgroups (P 0.05). None of the LLMs could identify any mpox correctly. Conclusion LLMs showed limited accuracy for diagnosing anogenital conditions, particularly for challenging images. The best-performing model achieved only 39.0% for top-1 diagnosis, indicating that current LLMs cannot reliably diagnose anogenital conditions. These tools may support supervised clinical triage but need further validation before routine clinical use.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Soe et al. (2026) studied this question.

synapsesocial.com/papers/6a0172233a9f334c2827244ahttps://doi.org/10.1093/infdis/jiag258
Ask AI
Helpful
Bookmark
Share
View Full Paper