Methodological analysis reveals distinct mathematical properties between resampling and perturbation fragility in trial tables, indicating the fragility index is not redundant with p-values.
The charge that the fragility index is a P-value in disguise rests on a strong correlation between the two across trials, and it overlooks a distinction between two distinct constructs derived from the same observed result. Resampling fragility asks how often the significance verdict would change if the trial were drawn again at the same size; it is a probability, and because it is based on the same observed result and a specified replication model, it is often strongly associated with the P-value. Perturbation fragility asks how many recorded outcomes must change before the observed table crosses the significance threshold; it is a count, the global fragility index, or a proportion, the global fragility quotient, and it measures the geometric distance of the observed table from the decision boundary, a quantity the P-value does not encode. Worked examples from published trials show tables with similar P-values and widely different perturbation fragility. The redundancy critique, whether made by resampling or by machine-learning models that use the P-value as a predictor, is correct about resampling fragility but does not address the fragility index.
No takes yet. Share an insight, caveat, or question.
Thomas F Heston (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: