Recently, we proposed two neutrality tests (Depaulis and Veuille 1998 ) based on haplotype number (K) and haplotype diversity (H). They relied on coalescent simulations conditional on the observed number of segregating sites (S) following the coalescent simulation procedure proposed by Hudson (1993) . In a companion letter, Markovtsova, Marjoram, and Tavaré (2001) use an alternative approach, based on the joint distribution of K and S, and show that the corresponding tests are not independent of the population mutational parameter θ (θ = 4Neμ, where Ne is the effective population size and μ is the neutral mutation rate per generation). They use the classical procedure of coalescent simulations conditional on θ and restrict their distribution to the particular subset of genealogies consistent with a particular value of S. They show that if θ is extreme, the probability of rejection can be substantially different from 5%. In another companion letter, Wall and Hudson (2001) show that the test based on K is reasonably robust in its original form. They perform coalescent simulations conditional on θ for a wide range of values. In contrast to the previous approach, they consider all the outcomes of the neutral simulations with various S values and look at the corresponding (various) confidence intervals given by our (Depaulis and Veuille 1998 ) procedure regardless of the θ value in the input of their simulations. This latter approach may be a better representation of a neutral distribution of genealogies. They study various statistics, including K, and the resulting type I error given by the confidence interval of Depaulis and Veuille (1998) remains close to 5% once corrected for the discreteness of the statistics. In practice, the exact value of θ is unknown, and we generally have no information on its value independent of a given data set. One should find a reliable procedure to account for uncertainty on θ. Replacing θ with an estimate would not be conservative. Rather than conditioning on this unknown parameter, we chose to condition directly on the observed value of S. In the present letter, we question the relevance of considering extreme θ values in addressing the robustness of the tests. We first show that those values of θ that lead to nonrobust neutrality tests with our procedure are highly unlikely given S under a neutral model. Second, we show with a Bayesian approach that our tests conditional on S are reliable, thus confirming Wall and Hudson's (2001) simulation results by an alternative approach. All values of θ are not equally likely given an observed value of S under the neutral model. Using Hudson's (1990) recursion, we computed the probability of obtaining an S value equal to or more extreme than a given value (the parameter values used by Markovtsova, Marjoram, and Tavaré [2001, tables 1 and 2] ). Only when θ = 10 is the S value not highly unexpected (table 1 ). For this θ value, the tests are conservative according to Markovtsova, Marjoram, and Tavaré (2001, table 1 ). We also computed Watterson's estimate of θ given S and its confidence interval following Kreitman and Hudson's (1991) method. The 95% confidence interval for θ always shows a much smaller range (1.3–27) than the 1–100 range used by Markovtsova, Marjoram, and Tavaré (2001). As pointed out by Wall and Hudson (2001), the fact that the observed S value is highly unexpected given θ is a sufficient reason to reject the null Wright-Fisher neutral model, and there is no need to use any other neutrality test. The reliability of the test should be assessed within the confidence interval of θ. To do this, we used the rejection algorithm suggested by Markovtsova, Marjoram, and Tavaré (2001) for θ values at the bounds of its confidence interval given S (table 1 ). We found a good overall fit between the nominal value of the test and its frequency of rejection (<12% in any case). However, note that θ values should be weighted by their probabilities given the data. Indeed, a difficulty with the procedure used by Markovtsova, Marjoram, and Tavaré is that the values for θ were taken arbitrarily. The probability of obtaining the configuration of values used in each simulation is thus ignored and we can hardly draw firm conclusions. Markovtsova, Marjoram, and Tavaré (2001) can only conclude that the confidence interval tends to narrow “as θ tends to zero or infinity.” Finally, as noted by Depaulis and Veuille (1998) and Wall and Hudson (2001) , the distribution of haplotypes depends to a large degree on recombination, and the tests should be used with caution if recombination is not zero. The extreme values of θ used by Markovtsova, Marjoram, and Tavaré (2001) would be even more unlikely under a model that includes recombination, since recombination tends to decrease the stochastic variance of estimates of θ (Hudson 1983 ). Yun-Xin Fu, Reviewing Editor Keywords: coalescent theory simulations neutrality tests haplotype distribution Address for correspondence and reprints: Frantz Depaulis, Institute of Cell, Animal and Population Biology, Ashworth Laboratory, King's Buildings, West Mains Road, Edinburgh EH9 3JT, United Kingdom. frantz.depaulis@ed.ac.uk. Table 1 Rejection Probabilities of the Haplotype Tests and Probability of S Given Various {θ} Values Table 1 Rejection Probabilities of the Haplotype Tests and Probability of S Given Various {θ} Values Fig. 1.—Posterior distribution of θ given S, obtained from equation (2) and assuming a uniform prior distribution of θ between 0 and 100. Parameters are identical to those in table 1 : s = 10, n = 10 (dashed curve); s = 40, n = 20 (thin solid curve); s = 50, n = 50 (dotted curve); s = 44, n = 20 (bold solid curve) We thank N. Barton, M. Cobb, Y. X. Fu, I. Gordo, A. Navarro, and S. Otto for helpful discussions and comments on earlier versions of this manuscript, and S. Tavaré for providing Markovtsova, Marjoram, and Tavaré's (2001) manuscript via his website. A computer program that implements Markovtsova, Marjoram, and Tavaré's (2001) rejection algorithm and the H and K haplotype tests conditional on either S or θ and on a value of the population recombination parameter are available from smousset@snv.jussieu.fr. F.D. was supported by NERC and S.M. and M.V. were supported by Groupe de Recherche GDR 1928 of the Centre National de la Recherche Scientifique.
No takes yet. Share an insight, caveat, or question.
Depaulis et al. (2001) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: