Panagiotou and Ioannidis1 (PI) have examined what they term ‘borderline associations’, based on P-value thresholds, and conclude that current genome-wide significance (GWS) guidelines (e.g. 5 × 10−8) are too stringent. To remedy this problem, PI recommend a P-value threshold of 10−7, though the statistical rationale for this particular figure is not clear. I agree with PI that the current criteria for declaring an association as ‘significant’ are often not appropriate and will here lay out the arguments for a Bayesian remedy. In general, a major difficulty with P-values is on deciding on a threshold for significance and this is true even for a single test. The current norm within the genome-wide association study (GWAS) literature is for P-value thresholds to be based upon controlling the family-wise error rate (FWER) (i.e. the probability of a single incorrect null rejection) at a very low level, typically 0.05. For example, with m = 1 000 000 tests the Bonferroni approach to controlling the FWER at 0.05 gives a threshold of 0.05/1 000 000 = 5 × 10−8. This approach, I will argue, is ill-advised since it pays no regard to the power of the tests. In particular, the same threshold is suggested for all sample sizes. A more appealing approach is to have a threshold that changes with power. In this way, a procedure can be formulated that produces both type I and type II errors that decrease to zero as hypothetical studies of increasing sample size are conducted. PI highlight various approaches to the generic problem of determining significance but do not discuss Bayesian approaches. Below, a Bayesian approach for testing many hypotheses is described which clearly illustrates the role of power. Put simply, all P-values are not born equal because knowledge of the associated power is vital for interpretation. PI have carried out a tremendous amount of work in cataloging associations that are borderline significant and in tracking down those associations to see if they are reproducible. Unfortunately, however, the power associated with the analysis of each SNP (which may be calculated from the sample size and minor allele frequency) is not reported. The approach we detail can either be used within a fully Bayesian approach, or for those not willing to fully commit, can be viewed as a mechanism by which a ‘sensible’ threshold for significance can be determined. 2 × 2 times table of costs when there are two hypotheses; CII is the cost of a type II error and CI the cost of a type I error 2 × 2 times table of costs when there are two hypotheses; CII is the cost of a type II error and CI the cost of a type I error In Table 2 below, we assume R = 1 so that type I and type II errors are equally costly. The table shows the P-value that corresponds to the Z2 threshold for a range of sample sizes and prior probabilities π0 = Pr(H0). Notice first how the P-value thresholds go to zero as the sample size increases. For π0 = 0.5 (PO = 1) and n = 20, 50, 100, the thresholds are ~0.05. It is interesting that these values are in the ballpark of the sample sizes that Fisher (who is accredited to popularizing the 0.05 threshold) would have been working with and presumably in the experiments in which he was involved the prior on the null would not have been close to 1 or 0. For example, in Tables 29 and 30 of Statistical Methods for Research Workers10 the sample sizes were 30 and 17 and Fisher discusses the 0.05 limit in each case, though in both cases he concentrates more on the context than on the absolute value of 0.05. P-value thresholds when costs of type I and type II errors are equal, as a function of the prior on the null, π0 and the sample size n The figures in bold represent situations in which the 0.05 threshold is approximately appropriate. P-value thresholds when costs of type I and type II errors are equal, as a function of the prior on the null, π0 and the sample size n The figures in bold represent situations in which the 0.05 threshold is approximately appropriate. To conclude: the 0.05 P-value threshold can be justified in some situations, but it is not reasonable to use this value as a universal rule. The GWAS situation is far different from that considered above because the prior on the null is so much closer to 1, and the sample sizes vary over a larger range. If the prior odds of the null, PO, increases, the threshold increases corresponding to a more stringent rule. If the relative cost of type II to type I errors, R, increases, the threshold decreases to give a more liberal procedure. Beyond a certain point, as n increases the type I error decreases to zero. The threshold depends crucially on the sample size, but not on the number of tests being performed. In contrast, frequentist procedures depend on the number of tests, but not on the sample size. We now turn to the thorny issue of how to decide upon a threshold. If one has a good prior estimate of π0 and is willing to specify the ratio of costs R, then one can proceed directly using equation (5) (the Bayes factor is relatively insensitive to the choice of prior variance W). A more pragmatic approach is to treat the Bayes factor as a device by which the ‘correct’ behaviour of the threshold can be obtained as a function of sample size and MAF. To illustrate, and to compare with Bonferroni, consider the situation in which we carry out a case–control study with n = n1 cases and n = n1 controls. We use the form (2) and assume a MAF of 0.5. Further suppose we have m = 1 000 000 tests and we set the ratio of costs at 10 (so that type II errors are 10 times as costly as type I errors) and π0 = 1 − 1/100 000. For the prior on the effect size suppose that there is a 95% chance that the log odds ratio lies between [− log 2, + log 2] to give W = 0.422. For comparison, we base the Bonferroni threshold on controlling the FWER at 0.05 so that the P-value threshold is 0.05/1 000 000 = 5 × 10−8. In Figure 1A we plot the type I error for both of the procedures; we see that the type I error is constant under Bonferroni by construction while the Bayes type I error is decreasing as a function of sample size. The Bonferroni procedure with an FWER of 0.05 is highly conservative and so the power is reduced, as is demonstrated in Figure 1B which shows the type II error as a function of sample size. Type I and type II errors can be difficult to translate into practical implications. The expected number of false discoveries (EFD) is m0α, where m0 is the the true number of null associations out of m and α is the type I error. Similarly, the expected number of true discoveries (ETD) is m1(1 − β) where m1 is the number of true signals. In practice, m0 (the number of null associations) and m1 are unknown but as an illustration suppose there are 50 true signals m1 = 50 out of m = 1 000 000. In Figure 1C we plot the EFD versus sample size and see that the Bayes procedure starts with around two false discoveries and then decreases as n increases. Bonferroni has an EFD of 0.04999 for all sample sizes. The benefit of the increased number of false discoveries under the Bayes procedure is the increased number of expected true discoveries as illustrated in Figure 1D. For example, for n = 3000 the expected number of true discoveries for Bayes is 27.9 and for Bonferroni is 16.2. Hence, in this example, trading around 2 false discoveries for 10 extra finds would seem beneficial. These sorts of simulation experiments can be carried out before analysis with the required operating characteristics being examined by altering π0 and R. Operating characteristics of Bonferroni and the threshold rule based on the Bayes factor. Type I and type II errors are displayed in the top row, and the expected number of false discoveries (EFD) and expected number of true discoveries (ETD) on the bottom row. This simulation is based on a situation in which the total number of tests is m = 1 000 000 and the true number of associations is m1 = 50 Type I error as a function of sample size for Bonferroni and for the rule in which we fixed the type I error at 5 × 10−8 at n⋆ = 3000 (which is indicated by the vertical line) We finally note that more sophisticated Bayesian methods for meta-analysis in a GWAS context have recently been described (Wen and Stephens).11 PI state that ‘The GWS should account for the multiplicity of comparisons’, but as we have seen, this is not the case in a Bayesian approach. The Bayes threshold boundary depends crucially on the sample size n but not on the number of tests m; the usual frequentist boundaries depend on m but not on n. We have concentrated on emphasizing that boundaries should depend on sample size but the power also depends on the MAF. For MAFs in the range 0.1 to 0.5 the power does not change too greatly. However, in the future it is likely that reliable data on SNPs with low MAF will be obtained and in this case the implications for a threshold will be more marked. In these cases, it would be beneficial to allow the the variance of the prior W to depend on MAF. For example, we might anticipate larger effect sizes at small MAF. The information contained in Table 1 of PI is helpful in this regard as it gives both the MAF and the estimated effect size, see also Park et al.12 Of course, when collecting together results on effect sizes and MAFs, in order to specify a form for W that depends on the MAF, one must account for the fact that we are unlikely to be seeing effects at low MAFs because of low power. In other words, the selection bias of the signals we are seeing must be considered. Software to carry out the calculations described in this article, written in the R language is available at: http://faculty.washington.edu/jonno/cv.html This work was funded by the National Institutes of Health (grant NIH U01 HG 005157). None declared.
No takes yet. Share an insight, caveat, or question.
Jon Wakefield (2012) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: