PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 2025Oral Diseases2 citations

Accuracy of ChatGPT‐4 Plus in Providing Information on Oral Cancer Management

View Full Paper
BYBüşra YılmazEKEmine Nur KahramanMBMichael T. Brennan

Key Points

  • ChatGPT-4 Plus demonstrated an accuracy rate of 63% for oral cancer management questions, emphasizing effective recovery guidance.
  • The highest accuracy was found in recovery (72%), while diagnosis responses were notably inconsistent, with only 55% rated as accurate.
  • Interrater reliability among specialists was strong, with ICC scores ranging between 0.85 and 0.93, indicating consistent evaluation standards.
  • Results indicate the potential for using AI in clinical settings, but advocate for clinician supervision to ensure patient safety.

Abstract

ABSTRACT Objective Artificial intelligence (AI)‐driven large language models, such as Chat Generative Pre‐Trained Transformer (ChatGPT)‐4 Plus, are increasingly used for patient education and clinical decision support in oral oncology, although their accuracy in oral cancer (OC) management remains uncertain. This study evaluates the accuracy of ChatGPT‐4 Plus responses to clinically relevant questions regarding OC diagnosis, treatment, recovery, and prevention. Methods A cross‐sectional study assessed 65 clinically relevant OC‐related questions using a paid ChatGPT‐4 Plus subscription without modifications. Three oral medicine specialists and one radiation oncologist rated accuracy on a four‐point scoring system. Interrater reliability was measured with the intraclass correlation coefficient (ICC), and chi‐square tests were used for comparisons. Results Among 65 questions, 63% of responses were Score 1, with none rated as Score 4. Score 1 was most frequent in Recovery (72%), followed by Treatment (62%), Prevention (60%), and Diagnosis (55%). Scores 2 and 3 responses were highest in Diagnosis (45%). Recovery had significantly higher Score 1 responses than Diagnosis ( p < 0.05), while other comparisons were not significant. ICC ranged from 0.85 to 0.93. Conclusions ChatGPT‐4 Plus provided accurate responses to clinically relevant OC‐related questions, particularly regarding recovery. However, diagnostic inconsistencies highlight the need for clinician oversight before integrating AI into practice.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yılmaz et al. (2025) studied this question.

synapsesocial.com/papers/68e24e60d6d66a53c24732b4https://doi.org/10.1111/odi.70110
Ask AI
Helpful
Bookmark
Share
View Full Paper