Randomized trial compares AI and radiologist readings for diagnosing prostate cancer, indicating potential benefits and limitations.
Background Artificial intelligence (AI) is increasingly used in prostate cancer diagnostic workflows but remains insufficiently validated in high-prevalence cohorts often encountered in academic referral centers. Purpose To compare the diagnostic performance of licensed AI software with routine radiologist readings of prostate MRI, using histopathology as the reference standard. Material and Methods In this retrospective study, 1000 patients underwent prostate MRI for suspected prostate cancer (between May 2020 and December 2024), followed by transperineal biopsy; 391 subsequently underwent radical prostatectomy. Diagnostic performance was assessed using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy across PI-RADS thresholds. Receiver operating characteristic (ROC) analysis, Cohen's kappa for inter-rater agreement, and paired McNemar's test were performed. Results In total, 959 patients (mean age = 69.7 ± 8.4 years) were included; clinically significant prostate cancer (csPCa) was found in 830 (86.5%) patients. AI assigned more cases to PI-RADS 1–2 and fewer to PI-RADS 3 (κ = 0.388; P <.001). At the PI-RADS ≥3 threshold, radiologists showed sensitivity 96.0%, specificity 20.9%, accuracy 85.9%, PPV 88.7%, and NPV 45.0%, compared with 91.3%, 37.2%, 84.0%, 90.3%, and 40.0% for AI ( P <.001). ROC performance was comparable (AUC 0.763 vs. 0.739; P = .28). In the prostatectomy subgroup, AI demonstrated a higher false-negative rate (8.4% vs. 4.1%; odds ratio = 3.83, 95% confidence interval [CI] = 1.52–11.51). Conclusion AI showed diagnostic performance comparable to that of radiologists for csPCa detection, with no significant difference in AUC. AI reduced indeterminate PI-RADS 3 scores but missed more significant cancers.
No takes yet. Share an insight, caveat, or question.
Ferreira et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: