PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 14, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Machine learning-based prediction models for noninvasive respiratory support failure in acute respiratory failure: a systematic review and meta-analysis

KAKenan AkgünHAHajed M. Al-OtaibiGCGaspar R. Chiappa

Key Points

  • The study aims to assess the effectiveness of machine learning models in predicting noninvasive respiratory support failure in adults with acute respiratory failure.
  • Conducted a systematic review and meta-analysis following PRISMA 2020 guidelines.
  • Searched databases including PubMed, Web of Science, and Scopus from January 2010 onward.
  • Included cohort studies that developed or validated machine learning models for predicting NIRS failure.
  • Discrimination assessed using the area under the receiver operating characteristic curve (AUC).
  • Bias and evidence certainty evaluated using PROBAST-AI and GRADE.
  • Fourteen cohort studies with 34,500 patients were analyzed.
  • The pooled AUC was found to be 0.84 with significant heterogeneity (I 2 = 99.5%).
  • No significant differences were observed by validation strategy or type of support.
  • All included studies had a high risk of bias, leading to very low evidence certainty.

Abstract

Background Early identification of noninvasive respiratory support (NIRS) failure in acute respiratory failure (ARF) is clinically relevant, as delayed intubation is associated with worse outcomes. Machine learning-based prediction models have been proposed to support escalation decisions, but their performance and reliability remain uncertain. Objective To systematically evaluate the discriminative performance of machine learning-based models for predicting NIRS failure in adults with ARF. Methods We conducted a systematic review and meta-analysis following PRISMA 2020 guidelines and registered the protocol in PROSPERO (CRD420251167330). PubMed, Web of Science, and Scopus were searched from January 2010 to the final search date. Cohort studies developing or validating machine learning models to predict NIRS failure, primarily defined as endotracheal intubation, were included. Discrimination was assessed using the area under the receiver operating characteristic curve (AUC). Logit-transformed AUCs were synthesized using random-effects models with restricted maximum likelihood estimation and Hartung–Knapp confidence intervals. Risk of bias and certainty of evidence were assessed using PROBAST-AI and GRADE, respectively. Results Fourteen cohort studies comprising 34,500 patients were included. The descriptive pooled AUC was 0.84 (95% CI, 0.78–0.89) with extreme heterogeneity (I 2 = 99.5%) and wide prediction intervals. Subgroup analyses showed no statistically significant differences by validation strategy or type of noninvasive respiratory support. All studies were rated at high risk of bias, and the certainty of evidence was very low. Conclusion Machine learning-based models demonstrate moderate discrimination; however, extreme heterogeneity, high risk of bias, and very low certainty of evidence preclude clinical implementation. Systematic review registration https://www.crd.york.ac.uk/PROSPERO/view/CRD420251167330 .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Akgün et al. (2026) studied this question.

synapsesocial.com/papers/69ddd8eee195c95cdefd66a4https://doi.org/10.3389/fmed.2026.1775670
Ask AI
Helpful
Bookmark
Share
View Full Paper