Key points are not available for this paper at this time.
Dear Editor, Statistical testing is a highly misunderstood topic in clinical research.1 The "significant vs. non-significant" dichotomy concerning the sole null hypothesis of no effect often leads to overconfident public health decisions. Interpreting confidence intervals (CIs) as ranges of compatibility can mitigate this problem. Indeed, the P value is a continuous index of the data's compatibility with the target hypothesis as assessed by the chosen statistical model (whose background assumptions must be true).2,3 A P-value close to 1 (respectively, 0) indicates high (respectively, low) compatibility. For instance, a hazard ratio (HR) = 3.5, 94% CI = (1.0, 12.2) tells us that all hypotheses contained within that interval (e.g., HR* = 1.1 or HR* = 12.1), according to the adopted statistical model, have P-values > 0.06 (since (1 − 0.06) × 100 = 94%). Assuming the statistical model is valid, the null hypothesis HR* = 1.0 has the same compatibility with the experimental result as the markedly non-null hypothesis HR* = 12.2 (P = 0.06, as these are the interval limits). Furthermore, the hypothesis most consistent with the data is not the null one, but HR* = 3.5 (P value = 1). Thus, the whole scenario – usually and mistakenly considered as "non-significant" only because the null P-value > 0.05 – merely suggests high statistical uncertainty (if the model is valid) and not a negligible effect size (the best point estimate of 3.5 is consonant with a pronounced effect size). Confidence (or, better, compatibility) intervals could be large due to violations of statistical background assumptions (e.g., multicollinearity), human errors or biases, or poor data fitting. However, it should not convey the message that narrow intervals are always indicative of high precision, since such property is still conditional on the same assumption of global validity.2 Hence, only if all underlying scientific and epistemological hypotheses are true (a utopian scenario), a small compatibility interval signals low statistical uncertainty (but not scientific uncertainty). Consequently, statistical inference necessitates numerous concordant studies. Moreover, health consequential decisions must be based on practical risks, costs, and benefits, that is, considering various evidence (e.g., biological, clinical, ethical, etc.) in addition to statistics. This message is of utmost importance in all sectors where the consequences for stakeholders can be substantial. The authors of this letter suggest adapting research standards based on the most recent and consolidated evidence on statistical testing1-5: Check all underlying assumptions and uncertainty sources, and report the assessment procedures in a supplementary file. Check the compatibility of the data with hypotheses of low and high effect sizes (e.g. HR = 1.0, HR = 1.1, and HR = 1.2 vs. HR = 3, HR = 4, and HR = 5). Assuming that point 1 has been appropriately performed, if the data are highly compatible (e.g., P > 0.20) with both of these types of effect sizes, then there is only strong statistical uncertainty (and not the absence of significance). Formulate conclusions based on the type of research. Terminal conclusions (e.g., recommendations for policymakers) must always be informed by decision analysis. Financial support and sponsorship Nil. Conflicts of interest There are no conflicts of interest.
Rovetta et al. (Tue,) studied this question.