PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 2026Baylor University Medical Center Proceedings0 citations

Evaluating the quality of artificial intelligence responses to psoriasis-related clinical and patient questions: a comparative study of ChatGPT, Gemini, and Microsoft Copilot

View Full Paper
GDGözde Ulutaş DemirbaşEDEsin DiremsizoğluADAbdullah Demirbaş

Key Points

  • This study aims to assess the quality of AI chatbot responses to psoriasis-related questions across different domains.
  • Fifty-four psoriasis-related questions were asked to ChatGPT, Gemini, and Microsoft Copilot.
  • Responses were scored by three board-certified dermatologists based on accuracy, completeness, and safety.
  • Scores for different categories included diagnostic, treatment, and patient questions.
  • Gemini scored the highest in treatment questions (8.00 ± 0.00) and significantly outperformed others in patient questions (P < 0.001).
  • Interrater agreement was substantial, with ChatGPT scoring κ = 0.743, Gemini κ = 0.830, and Copilot κ = 0.844.
  • All models showed high clinical safety, but responses varied in completeness and accuracy.

Abstract

Background Patient use of artificial intelligence (AI) chatbots for dermatologic information is increasing, but their performance on psoriasis-related questions across clinically distinct domains remains unclear. We compared ChatGPT (GPT-5.3 Instant), Gemini (Gemini 3 Flash), and Microsoft Copilot using a multidimensional scoring framework.Methods Fifty-four psoriasis-related questions were submitted to each model across diagnostic (n = 12), treatment (n = 12), and patient-question (n = 30) categories. Three board-certified dermatologists independently scored responses for accuracy, evidence consistency, completeness, and clinical safety (maximum score, 8).Results Interrater agreement was substantial to almost perfect (κ = 0.743 for ChatGPT, 0.830 for Gemini, and 0.844 for Copilot). Overall mean scores differed significantly: 7.36 ± 0.97, 7.77 ± 0.77, and 7.05 ± 1.09, respectively (Friedman P < 0.001). No difference was observed for diagnostic questions (P = 0.37). Gemini outperformed both models in treatment (8.00 ± 0.00) and patient questions (P < 0.001), while ChatGPT and Copilot did not differ in treatment. Differences were driven by completeness and accuracy, not clinical safety or evidence consistency. Gemini also had the highest rate of high-reliability responses (92.6%).Conclusions All models showed high clinical safety, but Gemini provided the most complete and highest-quality responses. The observation that models may provide accurate yet clinically incomplete responses, particularly for treatment content, emphasizes the need for physician oversight when AI-generated information is used in dermatological practice.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Demirbaş et al. (2026) studied this question.

synapsesocial.com/papers/6a2ba4778101cf8926f02ee9https://doi.org/10.1080/08998280.2026.2683945
Ask AI
Helpful
Bookmark
Share
View Full Paper