AI-generated echocardiographic segmentations were preferred over manual ones in 50.4% vs 36.6% of cases, indicating AI's potential to reduce interobserver variability.
Does AI-generated echocardiographic segmentation improve clinical acceptance compared to expert manual segmentation?
AI-generated echocardiographic segmentations are preferred by experts over manual segmentations, highlighting their potential to standardize measurements and reduce interobserver variability.
Absolute Event Rate: 0% vs 0%
Abstract Introduction Artificial intelligence (AI) segmentation models are increasingly adopted in healthcare, as they alleviate the significant time and effort required for manual delineation 1. However, while the performance of these algorithms is predominantly assessed using analytical metrics (e.g., Dice score, Hausdorff distance), their clinical applicability is less frequently investigated. Purpose This study aims to evaluate the clinical acceptance of AI-generated echocardiographic segmentations in comparison with expert manual segmentations. Methods A sample of 780 studies (patient mean age 64.1 ± 13.2 years; 59% male) was used to train an AI model based on the state-of-the-art nnUNet framework 2. In this cohort, cardiac structures (myocardium, ventricles, atria, and aorta) were delineated in apical 4-chamber (A4C), apical 2-chamber (A2C), parasternal short axis (PSAX), and parasternal long axis (PLAX) views at end-systole and end-diastole by 22 expert sonographers (mean expertise 5.17 ± 3.13 years). A testing subset, consisting of 4893 frames (1322 A4C, 1097 A2C, 1321 PSAX, and 1153 PLAX), was used to assess AI model performance, with cardiac structures inferred in these frames for which manual segmentation was available. This subset was entirely unseen by the model and therefore excluded from training. Five multi-centre experts from the pool of sonographers performed a blinded pairwise comparison of manual and AI-generated segmentations for each testing frame, selecting the segmentation deemed most suitable for clinical use. If neither segmentation met acceptable standards, experts could discard both. Results The inherent inter-observer variability in echocardiography was evidenced by a 13.0% rejection rate (636 images), where both AI-generated and manual segmentations were deemed unsuitable for clinical use by expert peer evaluators (Figure 1). AI-generated segmentations were preferred over manual segmentations: 2468 (50.4%) vs 1789 (36.6%) images. Indistinguishability, i.e., an even preference rate between AI-generated and manual segmentations, could be considered an optimistic result given the supervised nature of the AI model. The even superior performance of the AI model may be attributed to its ability to generalize across multiple experts, thereby reducing interobserver variability and resulting in consensus segmentations. This trend was maintained across views, except in the PLAX view. The underperformance of the AI model in PLAX may be explained by the lack of consensus in this view, as indicated by the high rejection rate (20%), which may have confused the model. Conclusions This study evidences the clinical acceptance of AI-generated segmentations in echocardiography. The superior performance of AI highlights its potential to standardize segmentations and reduce interobserver variability, without compromising clinical value. Ultimately, this work supports the integration of AI algorithms into medical practice.
Garcia-Sineriz et al. (Sat,) reported a other. AI-generated echocardiographic segmentations were preferred over manual ones in 50.4% vs 36.6% of cases, indicating AI's potential to reduce interobserver variability.