Key result
All four risk scores showed high discrimination for 30-day mortality, with SORT exhibiting the best discrimination (AUROC 0.922) and calibration, though all scores consistently over-predicted risk.
Why the study?
Surgical risk prediction tools should be externally validated in target populations prior to implementation to facilitate shared-decision-making and efficient allocation of perioperative resources.
Do surgical risk scores (SORT, NZRISK, POSSUM, P-POSSUM) accurately predict 30-day mortality in a general surgical population?
Population
44,031 surgical patients (53,395 operations) at Royal Perth Hospital from 2014 to 2021
Comparison
SORT, NZRISK, POSSUM, and P-POSSUM risk prediction scores
Design
Retrospective external validation cohort study
Follow-up
30 days
Authors
Loading...
SORT may be preferred for 30-day surgical mortality prediction; leaves open need for local recalibration before routine use.
Observational (n=44,031)
No
Do surgical risk scores (SORT, NZRISK, POSSUM, P-POSSUM) accurately predict 30-day mortality in a general surgical population?
Effect estimate: AUROC 0.922 for SORT
The SORT risk score demonstrated the best external validity for predicting 30-day mortality after surgery, though all evaluated scores tended to over-predict absolute risk, limiting their utility for shared decision-making without local recalibration.
Torlot et al. (2022) conducted an observational in Surgical patients (n=44,031). Surgical Outcome Risk Tool (SORT) and other risk scores (NZRISK, POSSUM, P-POSSUM) was evaluated on Risk score discrimination of 30-day mortality evaluated by area-under-receiver operator characteristic curve (AUROC) (AUROC 0.922 for SORT). All four risk scores showed high discrimination for 30-day mortality, with SORT exhibiting the best discrimination (AUROC 0.922) and calibration, though all scores consistently over-predicted risk.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: