Emergency department (ED) triage systems are essential for prioritizing patients based on clinical urgency, yet traditional methods are subject to inter-observer variability and limited predictive accuracy. Artificial intelligence (AI) and machine learning (ML) have emerged as promising tools to enhance triage decision-making. This systematic review synthesizes existing evidence on AI/ML-based triage systems in EDs, focusing on predictive performance and reported clinical outcomes, while critically appraising methodological quality and gaps affecting clinical applicability. A comprehensive search was conducted across PubMed, Scopus, CINAHL, IEEE Xplore, and Web of Science (2021-2026), supplemented by citation tracking. Studies developing or validating AI/ML models for ED triage and reporting quantitative performance metrics were included. The Prediction Model Risk of Bias Assessment Tool (PROBAST) was used for quality assessment. A narrative synthesis was performed due to substantial heterogeneity. Fourteen retrospective observational studies (2021-2026) from eight countries met the inclusion criteria (sample sizes: 657 to >2.6 million visits). Models included gradient-boosted trees, random forest, logistic regression, neural networks, natural language processing, and large language models. AUC-ROC ranged from 0.642 to 0.991 (highest for mortality (0.874-0.933) and pediatric critical illness (0.991)). However, calibration was reported in only three studies, external validation in only five, and only one study demonstrated direct clinical process improvement (reduced missed ECGs). No prospective or randomized controlled trials were identified. PROBAST rated 11 studies as low risk of bias, two as high, and two as unclear. AI/ML models show moderate to excellent retrospective predictive performance for ED triage outcomes, particularly ensemble tree-based and NLP-enhanced approaches. However, the evidence base is severely limited by overreliance on heterogeneous retrospective designs, insufficient calibration reporting, and lack of prospective or external validation. Consequently, the strength of conclusions regarding clinical applicability remains weak. Future research must prioritize rigorous prospective validation, calibration reporting, and randomized trials measuring patient-centered outcomes before clinical implementation is considered.
Mohamed et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: