Key points are not available for this paper at this time.
False-positive results that arise as the result of chance are common in the medical literature 1-3. By chance, every study sample will have slight imbalances that don't reflect the whole population. If researchers look at enough characteristics of a given sample, they are bound to discover these quirks and conclude (mistakenly) that they have significance for the whole population. This is the problem of multiple testing—the more tests you run on a sample, the greater the likelihood of a chance finding. This article will formally describe the problem of multiple testing and give readers tools for spotting chance findings in the literature. In hypothesis testing, this is a false-positive error. The researcher concludes that an effect exists when it does not. The hypothesis of no effect—for example, the hypothesis that 2 variables are unrelated or that 2 groups don't differ. A type I error occurs when the null hypothesis is erroneously rejected. If 100 statistical tests are run when: (1) there are no real effects; and (2) these tests are independent, what is the probability of at least one false positive (that is, one P value under .05)? For 1 test: If there are no real effects, the probability of a false positive arising in a given test is .05. So, the probability that a false positive does not occur is 1 - .05 = .95. For 100 tests: If the tests are independent, meaning they are unrelated to each other, then the probability that no false positives occur in 100 tests is: .95100 = .006. Thus, the probability that at least one false positive does occur is 1 - .006 = .994, or 99.4%. Note that this calculation requires two key assumptions: 1. The null hypothesis is true for all 100 effects being tested. 2. The effects being tested are completely independent. In many cases where multiple tests are run in the literature, one or both of these assumptions may not be true. A measure of the magnitude of an observed effect—for example, how big the difference between groups is. A measure of relative risk formed by dividing the risk in one group by the risk in a reference group. Values of 1.0 indicate no difference in risk; values >1.0 indicate increased risk; and values 2 cm), lymph node metastasis (yes/no), and histologic grade (well/moderate/poor). Figure 2 shows the distribution of the resulting 50 P values. Four P values were less than .10, which is consistent with chance. (In addition to the 3 P values mentioned previously, decaffeinated coffee was linked to protection against breast cancer in postmenopausal never hormone users, P = .02; interestingly, the authors did not comment on this finding.) The effect sizes are also consistent with chance: risk ratios were close to the null value of 1.0 (ranging from 0.67 to 1.79), indicated protection (1.0), and showed no consistent dose–response pattern across increasing levels of consumption. It is easy to come up with a plausible sounding biological story to explain why caffeine is important in women with benign breast disease and for certain types of tumors, but, in fact, the most likely explanation for the findings is chance alone. The distribution of P values from 50 statistical tests examining the relationship between caffeine, coffee, and tea intakes and breast cancer 5. P values were taken from tests for trend only and from the most adjusted model when more than one model was presented (Tables 2-5 from Ishitani, et al 5). Four P values fall below .10, which would be expected due to chance. In judging whether a given finding in the literature is likely to be to the result of chance, readers should consider whether the analyses were hypothesis-driven or exploratory, how many tests were run, the size of the P values, the pattern of effect sizes, and whether P values were adjusted for multiple comparisons (Table 2). When evaluating the literature, readers should distinguish between hypothesis-driven analyses and hypothesis-generating or exploratory analyses. When researchers specify a priori (before the study is conducted) a small number of hypotheses that they are planning to test, including clear definitions of the predictors and outcomes, this is hypothesis-driven research. This approach limits the number of statistical tests run, thus controlling the overall type I error. In contrast, when researchers test a large number of hypotheses after the data have been collected—essentially searching through the data to find associations—this is hypothesis-generating, or exploratory, research, and the type I error rate is likely to be high. This is not to say mining the data in this manner is wrong or bad. Often researchers collect large amounts of data in the course of a focused study, and it would be a waste not to examine these data. But “statistically significant” results that come out of such an exploratory analysis should be regarded with a greater level of scrutiny. For example, if a randomized trial tests the effect of a drug versus a placebo in stroke patients, the P values associated with the primary hypothesis (the difference in recovery between drug-treated and placebo-treated patients) can be taken at face value. However, if in that same study, the researchers explored the associations between a large number of nutrition variables (that happened to be collected on a food frequency questionnaire as part of the study) and stroke recovery, then these results should be clearly identified as exploratory and interpreted cautiously. Readers can count the number of tests reported in a paper and multiply it by .05 to get a rough idea of the number of P values less than .05 that would be expected to arise by chance alone (if no effects being tested were real). Of course, the data presented in a paper usually represent a subset of all the statistical tests run (particularly for exploratory analyses), so keep in mind that the number of tests run may be much larger than the number of tests reported. There are a number of statistical approaches available that can reduce the number of tests being run. Global tests such as analysis of variance (ANOVA) and repeated-measures ANOVA compare multiple groups or multiple time points simultaneously, thus generating only one P value. For example, rather than running 3 t tests (and thus inflating the type I error) to compare the means in 3 groups, one can instead conduct a single ANOVA analysis that compares all three means at once. Statisticians also create composite outcomes to reduce the number of outcomes being tested. The strength of the evidence against the null hypothesis increases with smaller P values. Therefore, a P value of 1.0), but only a few results achieve moderate statistical significance, one may suspect that low statistical power is a factor rather than chance. On the other hand, if the effect sizes show no consistent pattern—as in the caffeine/breast cancer study mentioned previously—then chance may be a more likely explanation. Statisticians have devised ways to “adjust” P values or confidence intervals to account for the number of tests run. The basic idea is to preserve the overall type I error rate at .05 by lowering the threshold for statistical significance (to lower than <.05) or widening the confidence interval. For example, the simplest approach is the Bonferroni correction: When k tests are run, only P values under .05/k are deemed significant (eg, if 5 tests are run, only P values under .01 are reported as significant). Although easy to understand and conduct, the Bonferroni correction is overly conservative—it represents a “worst-case” scenario where all the tests being conducted are completely independent (which is usually not the case). Thus, many less conservative (but more mathematically intensive) methods have been developed. Applying these requires statistical software and/or consultation with a statistician. Formal corrections for multiple comparisons are most often used in the context of hypothesis-driven research. For example, if the authors plan to look at multiple outcomes, look at a limited number of planned subgroups, or engage in interim analyses, they may build in a correction for these multiple comparisons. For exploratory analyses, formal adjustment of P values (and confidence intervals) is usually impractical. In these contexts, it is difficult to precisely quantify the total number of tests run and their interrelatedness; and, because of the large number of tests run, the adjusted threshold for statistical significance may be so small that it may be unreachable (and the false-negative rate will be extremely high). For exploratory analyses, it is more important to judge P values cautiously than to try to formally determine their true significance level. Multiple testing is a major source of false positives in the medical literature. Exploratory analyses are particularly prone to this type of error and should be interpreted cautiously. When a few moderate size “significant” P values arise in the course of a large number of exploratory analyses, these likely reflect chance rather than real associations. Precise adjustment of P values and confidence intervals is often impractical in the context of exploratory research, but can be useful for hypothesis-driven research.
Kristin L. Sainani (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: