Abstract Rationale Pleural effusion (PE) is a common thoracic condition requiring accurate detection for optimal management. Thoracic ultrasound (TUS) is the first-line diagnostic modality due to superior sensitivity, portability, and safety, but remains operator-dependent. Artificial intelligence (AI)-based automated analysis may reduce variability and standardize interpretation. This systematic review evaluates the diagnostic accuracy of AI-assisted TUS for PE detection. Methods This systematic review was prospectively registered with PROSPERO (CRD420251128416) and conducted following PRISMA guidelines. A comprehensive literature search was performed across MEDLINE, Scopus, Google Scholar, IEEE Xplore, Cochrane CENTRAL, and ClinicalTrials.gov through 20 August 2025. Eligible studies evaluated AI-based automated analysis applied to TUS for PE detection, requiring recognized reference standards including expert interpretation or chest CT. Risk of bias was systematically assessed using QUADAS-2, and certainty of evidence evaluated with GRADE framework. Due to substantial heterogeneity in AI architectures, input data formats, reference standards, and clinical settings, structured narrative synthesis was performed instead of quantitative meta-analysis. Results Five studies enrolling 7,565 patients published between 2021-2025 were included. Four were retrospective and one prospective, conducted across China, Australia, Canada, and Republic of Korea. Clinical settings included emergency departments, intensive care units, inpatient wards, and ultrasound departments. All studies employed convolutional neural networks with diverse architectures: ResNet with Split-Attention, EfficientNet-B0, Attention U-net versus standard U-net, Regularized Spatial Transformer Networks, and custom Pleff-Net. Input formats included still images (three studies) or video sequences (two studies), with datasets ranging from 1,440 images to over 313,000 frames. Diagnostic performance showed sensitivity 70.6-100%, specificity 67-100%, overall accuracy 79.0-98.2%, and AUC 0.77-0.998. Most studies reported sensitivity and specificity exceeding 90% in internal validation. However, multicenter external validation demonstrated progressive attenuation: AUC declined from 0.94 (internal/temporal) to 0.89 (external test 1) and 0.77 (external test 2). Subgroup analyses revealed reduced accuracy for trace (77.0% sensitivity), small (92.2%), and complex effusions (77.0%), and in critically ill patients. Three studies implemented Grad-CAM visualization without formal clinical validation. QUADAS-2 identified high risk of bias in patient selection and index test conduct across all studies. Conclusions AI-assisted TUS demonstrates high diagnostic accuracy for PE detection in curated datasets but with considerable performance variability and methodological limitations. Certainty of evidence is moderate, constrained by retrospective designs, limited external validation, and lack of patient-centered outcomes. Prospective multicenter studies with rigorous external validation, standardized acquisition protocols, comprehensive explainability frameworks, and assessment of clinical impact are essential before widespread clinical implementation. This abstract is funded by: None
Marchi et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: