PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 26, 2026European Journal of Radiology Artificial Intelligence1 citationsOpen Access

Diagnostic Agreement Between ChatGPT and Nuclear Medicine Experts in the Interpretation of Cerebral 18FFDG-PET/CT for Neurodegenerative Disease

ASArndt‐Hendrik SchievelkampDUDominic UftonJHJan Heilinger

Key Points

  • This research aims to evaluate how accurately two AI models interpret neurodegenerative disease diagnostics compared to seasoned experts.
  • Analyzed 100 anonymized cases of neurodegenerative diseases through AI and expert interpretation.
  • AI models used textual descriptions of imaging findings and clinical history provided as input.
  • In a reproducibility test, 20 cases were reassessed for run-to-run agreement using an ordinal scale.
  • ChatGPT-4o and ChatGPT-5 had median agreement scores of 1.00 on a five-level scale.
  • Main diagnoses were correctly identified in 86% of cases (ChatGPT-4o) and 89% (ChatGPT-5).
  • Reproducibility rates showed 75% exact agreement for ChatGPT-4o and 55% for ChatGPT-5, with varied consistency levels.

Abstract

Background: Artificial intelligence is a valuable tool in medical imaging and diagnostics.This study assesses the agreement between ChatGPT-4o and ChatGPT-5, two large language models, and expert nuclear medicine physicians in interpreting 18 FFDG-PET/CT findings in neurodegenerative diseases.Methods: 100 anonymized cases were analyzed, comparing AI-generated differential diagnoses with expert reports.Models received only textual descriptions of imaging findings and clinical history extracted from the reports, with patient age provided as an additional input variable.In a 20-case reproducibility subset per model, we re-queried the same cases in a new chat session with the conversation history cleared, using the identical prompt, and assessed run-to-run agreement on a five-level ordinal scale (0, 0.25, 0.5, 0.75, 1).Results: Median agreement scores were 1.00 IQR 0.50-1.00with ChatGPT-4o and 1.00 0.75-1.00with ChatGPT-5.The main diagnosis was correctly identified in 86% (ChatGPT-4o) and 89% (ChatGPT-5) of cases, respectively.Both models performed best in cases with well-defined metabolic patterns but were less accurate for complex metabolic patterns associated with broad differentials.In the reproducibility subset, exact run-to-run agreement was 75% for ChatGPT-4o (quadratic-weighted = 0.48) and 55% for ChatGPT-5 ( = 0.65).The higher for ChatGPT-5 reflects greater consistency on the ordinal scale despite fewer exact matches, with most differences representing one-step shifts.Conclusions: While ChatGPT is not approved for clinical diagnostics, it demonstrated substantial diagnostic alignment with expert physicians when interpreting textual case information.Reproducible outputs highlight its potential as a supportive diagnostic tool.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schievelkamp et al. (2026) studied this question.

synapsesocial.com/papers/69edab424a46254e215b355chttps://doi.org/10.1016/j.ejrai.2026.100101
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Deep Semisupervised Transfer Learning for Fully Automated Whole-Body Tumor Quantification and Prognosis of Cancer on PET/CT2024 · 64 citations
  2. 2Reporting checklist for foundation and large language models in medical research (REFINE): an international consensus guideline2026 · 11 citations
  3. 3Applications of artificial intelligence and deep learning in molecular imaging and radiotherapy2020 · 200 citations
  4. 4Trustworthy Artificial Intelligence in Medical Imaging2021 · 92 citations
  5. 5Artificial Intelligence on FDG PET Images Identifies Mild Cognitive Impairment Patients with Neurodegenerative Disease2022 · 18 citations