Background: Artificial intelligence is a valuable tool in medical imaging and diagnostics.This study assesses the agreement between ChatGPT-4o and ChatGPT-5, two large language models, and expert nuclear medicine physicians in interpreting 18 FFDG-PET/CT findings in neurodegenerative diseases.Methods: 100 anonymized cases were analyzed, comparing AI-generated differential diagnoses with expert reports.Models received only textual descriptions of imaging findings and clinical history extracted from the reports, with patient age provided as an additional input variable.In a 20-case reproducibility subset per model, we re-queried the same cases in a new chat session with the conversation history cleared, using the identical prompt, and assessed run-to-run agreement on a five-level ordinal scale (0, 0.25, 0.5, 0.75, 1).Results: Median agreement scores were 1.00 IQR 0.50-1.00with ChatGPT-4o and 1.00 0.75-1.00with ChatGPT-5.The main diagnosis was correctly identified in 86% (ChatGPT-4o) and 89% (ChatGPT-5) of cases, respectively.Both models performed best in cases with well-defined metabolic patterns but were less accurate for complex metabolic patterns associated with broad differentials.In the reproducibility subset, exact run-to-run agreement was 75% for ChatGPT-4o (quadratic-weighted = 0.48) and 55% for ChatGPT-5 ( = 0.65).The higher for ChatGPT-5 reflects greater consistency on the ordinal scale despite fewer exact matches, with most differences representing one-step shifts.Conclusions: While ChatGPT is not approved for clinical diagnostics, it demonstrated substantial diagnostic alignment with expert physicians when interpreting textual case information.Reproducible outputs highlight its potential as a supportive diagnostic tool.
Schievelkamp et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: