PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026BMC Medical Informatics and Decision Making0 citationsOpen Access

Performance fairness of neural network models in early risk assessment of inpatients with varying severity: a retrospective study

LLLan LanSHShixin HuangYCYang Chen

Key Points

  • The study aims to evaluate the fairness of an LSTM model in predicting in-hospital mortality among inpatients with varying severity levels.
  • Utilized the MIMIC-IV database with records from over 50,000 ICU patients.
  • Divided patients into subgroups based on length of stay and SAPS II scores.
  • Trained an LSTM model on the training set and tested it across subgroups.
  • Employed metrics such as AUROC and AUPRC to assess model performance.
  • Conducted logistic regression and Bonferroni correction for performance comparison.
  • Overall AUROC was 0.834, with better performance for shorter LOS and lower SAPS II scores.
  • Highest AUROC of 0.931 observed in the [12, 94) hours LOS group.
  • AUROC of 0.811 found in the [0, 25) SAPS II score group.
  • Model accuracy decreased as LOS and SAPS II scores increased.
  • Logistic regression confirmed significant effects of LOS and SAPS II on model accuracy.

Abstract

To evaluate the performance fairness of a long short-term memory (LSTM) model in predicting in-hospital mortality for inpatients with varying severity, as reflected by length of stay (LOS) and initial clinical scores. This retrospective study used the Medical Information Mart for Intensive Care (MIMIC)-IV database, which includes records from over 50,000 ICU patients. Patients were divided into subgroups based on LOS and Simplified Acute Physiology Score (SAPS) II. The LSTM model was trained on the training set and then tested on these subgroups in the test set. Metrics such as area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), accuracy, sensitivity, and specificity were used to evaluate model performance. Statistical analyses, including logistic regression and Bonferroni correction, were conducted to compare performance across subgroups. The LSTM model’s performance varied significantly among different LOS and SAPS II score groups. The overall AUROC was 0.834, but the model performed better for patients with shorter LOS and lower SAPS II scores. The highest AUROC of 0.931 was observed in the [12, 94) hours LOS group, and 0.811 in the [0, 25) SAPS II score group. The model’s accuracy decreased with increasing LOS and SAPS II scores. Logistic regression confirmed that LOS and SAPS II scores significantly affected model accuracy, with longer LOS and higher SAPS II scores associated with poorer model performance. When using long-term outcomes like in-hospital death to build early assessment models, there are significant fairness issues in model performance across LOS and SAPS II groups. Developing dynamic prediction models using short-term outcomes may help reduce these fairness issues.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lan et al. (2026) studied this question.

synapsesocial.com/papers/69c8c371de0f0f753b39e45chttps://doi.org/10.1186/s12911-026-03448-7
Ask AI
Helpful
Bookmark
Share
View Full Paper