The Editor, I read with interest the prospective study by Gulia and colleagues, and commend the authors for the data collection effort and the attempt to follow TRIPOD. 1 Several methodological reflections, however, suggest the reported AUC of 0. 95 considerably overstates the model's true performance. 2, 3e AAGC Model Double-counts APACHE II APACHE II already contains age, creatinine, and a GCS-derived neurological score. Putting all four into one model is, in effect, fitting APACHE II twice. This explains the strong negative correlation the authors observed between APACHE II and GCS (r = -0. 76), and it makes the individual odds ratios uninterpretable. The model should use either APACHE II or its components, not both. 2 Categorization Throws Information AwaySplitting age at 60, grouping APACHE II into ordinal "risk categories, " and using qSOFA binary components alongside the continuous variables they were derived from (RR, SBP, GCS) reduces statistical power and biases estimates. 2, 4Continuous variables should stay continuous, with non-linearity handled by restricted cubic splines. 2 LASSO followed by Ordinary Refitting Undoes the ShrinkageThe authors selected variables with LASSO at λ₁se, then re-estimated coefficients without penalization. This discards the very shrinkage that makes LASSO useful out-of-sample, leaving inflated odds ratios and an over-confident nomogram. 5Either retain the penalized coefficients, or pre-specify a small set of predictors rather than letting the data choose.
Kapil Soni (Tue,) studied this question.