PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 2, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Exploring the Diagnostic Limits of Chatgpt: How Far Can a Large Language Model Go in Histopathological Image Interpretation?

View Full Paper
AMAghajan MusaliJMJamal Musayev

Key Points

  • This research evaluates the diagnostic performance of ChatGPT in interpreting histopathological images compared to experienced pathologists.
  • Evaluated 24 histopathological images by ChatGPT-4o mini and 15 experienced pathologists.
  • Standard diagnostic queries were used without clinical context.
  • Responses were classified as correct, false positive, false negative, low-impact error, or no interpretation.
  • Statistical analyses were performed using McNemar’s test and Fisher’s exact test.
  • ChatGPT-4o mini achieved 71.4% accuracy, 60.0% sensitivity, and 77.8% specificity.
  • Pathologists averaged 89.8% accuracy with 97.7% sensitivity and 87.1% specificity.
  • Low-impact errors were 33.3% for ChatGPT compared to 6.9% for pathologists.
  • Statistical analysis showed significant differences favoring pathologists.

Abstract

Aims: Artificial intelligence’s integration into pathology has accelerated with the adoption of digital workflows. Large language models like ChatGPT offer unique opportunities but have yet to be systematically evaluated in diagnostic image interpretation. Methods: In this comparative study, 24 histopathological images representing various tissue types and pathological entities were evaluated by ChatGPT-4o mini and 15 experienced pathologists. The model was prompted with a standard diagnostic query without access to clinical information. Pathologists independently assessed the same images. Responses were categorized as correct, false positive, false negative, low-impact error, or no interpretation. Standard diagnostic metrics were calculated, and group comparisons were conducted using McNemar’s test and Fisher’s exact test. Interobserver agreement among pathologists was analyzed using Fleiss’ kappa. Results: ChatGPT-4o mini achieved an accuracy of 71.4%, with a sensitivity of 60.0% and a specificity of 77.8%. The average accuracy of pathologists was 89.8%, with 97.7% sensitivity and 87.1% specificity. Low-impact errors were more frequent with ChatGPT-4o mini (33.3%) compared to pathologists (6.9%). McNemar’s test revealed a statistically significant difference in favor of pathologists. The interobserver agreement among pathologists was in the lower range. Conclusion: While ChatGPT-4o mini demonstrated partial diagnostic capabilities, it underperformed compared to experienced pathologists. The absence of a clinical context likely impacted the results. Future artificial intelligence models integrating image analysis and clinical data may enhance performance. Despite limitations, the potential ChatGPT holds as a supportive diagnostic tool in pathology is highlighted in this study.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Musali et al. (2026) studied this question.

synapsesocial.com/papers/69a528b3f1e85e5c73bf0452https://doi.org/10.4274/tmsj.galenos.2026.2025-10-1
Ask AI
Helpful
Bookmark
Share
View Full Paper