Optical coherence tomography (OCT) plays a crucial role in diagnosing retinal diseases, such as diabetic retinopathy (DR) and age-related macular degeneration (AMD), as well as in identifying neurodegenerative biomarkers. Despite advancements in U-Net-based convolutional networks for OCT image segmentation, there is a lack of systematic reviews comparing their performance with expert manual segmentations. This review aims to assess the efficacy of these automated networks in segmenting retinal fluid and pathology in OCT images. By searching three different databases, PubMed, Web of Science, and Scopus, over the past five years, we conducted this systematic review and meta-analysis using data from 16 diagnostic-accuracy studies. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool. The analysis used mean and standard deviation for the continuous outcomes and employed a random-effects model. Analyses were performed using Review Manager software version 5.4 (The Cochrane Collaboration, London, UK, 2020). Artificial intelligence (AI) and human Dice scores did not differ significantly (standardized mean difference (SMD) = -0.08; 95% CI: -1.16 to 0.99; p = 0.88), nor did intraclass correlation coefficient (ICC) values (SMD = -0.13; 95% CI: -5.70 to 5.45; p = 0.96). However, very high heterogeneity (I² > 90%) limits the reliability of these pooled estimates. AI achieved expert-level Dice scores for subretinal fluid (0.88-0.96) and geographic atrophy (0.94). Intraretinal fluid was more challenging (Dice 0.79-0.89). Volumetric reliability was strong (ICCs > 0.94). Device-dependent variability was substantial; kappa was 0.37 for ZEISS versus 0.73 for Spectralis, indicating a need for device-specific optimization. Volumetric analyses revealed minor systematic overestimation (mean difference: -0.05 mm²). Processing times ranged from 100 milliseconds per B-scan to several seconds per volume, representing substantial time savings versus manual segmentation. Fully automated U-Net pipelines reach expert-level accuracy for subretinal fluid and geographic atrophy but remain limited for intraretinal fluid and show marked device-dependent variability. Clinical translation requires four priorities: standardized multi-device benchmarks, domain adaptation for cross-platform robustness, hybrid AI-human workflows pairing automated pre-segmentation with expert oversight, and prospective clinical trials. These steps are needed to move AI segmentation from a research tool to a clinical decision-support system.
Awan et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: