The question of how DNA barcodes can and should be used in taxonomy has been debated for some time (Lipscomb et al. 2003; Tautz et al. 2003; Blaxter 2004; Vogler and Monaghan 2007; Wiens 2007). Although few doubt that they are a valuable molecular tool for matching unidentified specimens to described taxa, this has little to do with the question of whether barcodes can be used to delimit species in the first place. The most radical turn in this debate has been the plea for a DNA-based taxonomy (Tautz et al. 2003; Blaxter 2004; Pons et al. 2006; Vogler and Monaghan 2007). Its proponents argue that “the vast majority of sequence variation in nature is partitioned into clearly defined clusters” (Vogler and Monaghan 2007, p. 4), which “ […] broadly mirror the species category” (Papadopoulou et al. 2008, p. 1) and could thus serve as basic taxonomic units. Initial attempts to employ this “barcoding gap” have relied on defining cutoff values of sequence divergence a priori (e.g., Blaxter 2004). Considering that the amount of genetic diversity within species can vary by orders of magnitude, it is clear that such an approach is arbitrary at best. Pons et al. (2006) have recently proposed a likelihood method that circumvents this problem by testing for clustering in ultrametric trees. They argue that “these new quantitative approaches can infer the elusive species boundary directly from the transition in branching rate and constitute an exciting possibility to define species from sequence variation […]” (Vogler and Monaghan 2007, p. 6). Given such claims, it is not surprising that this method enjoys increasing popularity, having been applied to a number of mitochondrial DNA (mtDNA) data sets (e.g., Pons et al. 2006; Ahrens et al. 2007; Fontaneto et al. 2007; Papadopoulou et al. 2008). In essence, the Mixed-Yule-Coalescent model (MYC) of Pons et al. (2006) splices together the classical null models of macroevolution and microevolution. Unlike standard models of divergence that view the genealogical process as nested within the species tree, the MYC model assumes a single transition time T at which lineage sorting happens instantaneously and the branching of species clades is replaced by multiple independent coalescences occurring within them (Pons et al. 2006). Assuming T to be a particular node in the tree, Pons et al. (2006) use the internode intervals to find the maximum likelihood solution for T under the MYC model and compare this to the likelihood under a null model of a single neutral coalescent process. Although it has been pointed out that this and similar schemes relying on single locus data cannot deal with lineage sorting and thus necessarily fail to detect recently diverged lineages (Hudson and Coyne 2002; Pons et al. 2006), the potential problems arising from population structure have so far largely been ignored. In a recent paper, Papadopoulou et al. (2008) have tested the MYC method on genealogies simulated under a symmetric island model, which assumes a population divided into multiple demes or subpopulations that are connected to all other such demes through migration occurring at rate m (Wright 1931). Such population structure tends to produce clustering, very similar to that expected under the MYC model, simply because lineages residing in the same deme coalesce more rapidly on average than those in different demes (Fig. 1). In this setting, the genealogy of a sample may be related to the demic structure in 3 different ways: If gene flow is very low, clusters may correspond well to demes and may thus constitute meaningful taxonomic units in the broadest sense (leaving aside the question how species should be defined). If gene flow is high, clustering may be weak or nonexistent. Clusters may be essentially random, that is, only partially corresponding to demes. Genealogy of a sample taken from two demes B and C in an island model. Coalescence within demes happens rapidly compared with coalescence of lineages from different demes, which have to be preceded by migration events (dashed arrows). In this case, three “clusters” are produced because a lineage from B escapes within-deme coalescence through migration into an unsampled deme A and has to wait a long time until it finds itself in the same deme as the remaining lineage. In their simulations, Papadopoulou et al. (2008) assume an extreme sampling scheme where samples are taken from all demes. They find that clustering under the MYC model is only significant when migration rates are extremely low (Nm < 10−3) in which case clusters correspond very well to demes (Case 1) (Papadopoulou et al. 2008, figure 3). Once migration is above a certain threshold value, clustering disappears rapidly and is not detected by the MYC method (Case 2). The authors conclude that “the MYC approach appears to be conservative, only detecting the products of population isolation when the levels of gene flow are much lower than those traditionally regarded as sufficient for neutral population divergence” (Papadopoulou et al. 2008, p. 8). It is worthwhile to recall some basic properties of the coalescent for samples in an island population here. Going backwards in time, lineages have a probability m per generation of escaping coalescence in their local deme. The number of sampled demes over the total number of demes, d/D, is crucial in determining the fate of such escaping lineages. The first migrating lineage has probability d/D of landing in a sampled deme in which it may coalesce. Alternatively, with probability 1 − d/D, it lands in an unsampled deme and its coalescence has to be preceded by at least one additional migration event. Realizing the pivotal role of the sampling scheme, Wakeley (1998, 2008) has developed an elegant approximation for the coalescent in an island model. If the number of unsampled demes is large, d/D tends to zero and the ancestral process can be split into two phases occurring on different timescales (large D-approximation). Initially, lineages may either coalesce in their local deme or spread out into unsampled demes (scattering phase) (Fig.1). Once every lineage resides in a separate deme, the ancestral process is a neutral coalescent with a rate dependent on the total number of demes, D, their size, N, and m (collecting phase). This separation of timescales and the strong pattern of sequence clusters resulting from it may superficially resemble the two phases in the MYC model. However, there are two important differences. First, there is no branching process in the structured coalescent. Instead, the collecting phase is another, although much slower, neutral coalescent. Thus, theoretically, one could extend the likelihood approach of Pons et al. (2006) to distinguish between the two models. Second and more importantly, clusters in the structured coalescent may be the result of migration events into unsampled demes during the scattering phase and are thus fundamentally random (Case 3). One would therefore expect the sampling scheme to have a major impact on the performance of the MYC method. To investigate this, I repeated the simulations of Papadopoulou et al. (2008) for varying d/D. Genealogies were simulated in MS (Hudson 2002). The effect of the mutational variance on tree reconstruction was ignored, that is, the method was applied directly to simulated genealogies. Likelihoods under both the MYC and a single neutral coalescent were calculated using the genealogy package in Mathematica (available from www.biology.ed.ac.uk/research/institutes/ evolution/software/barton/index.html). For each replicate, the two models were compared in a likelihood ratio test and the number of inferred clusters recorded (Papadopoulou et al. 2008). To investigate the region of the parameter space for which the MYC method breaks down, the following sampling scheme was used. Genealogies were simulated for a total of 100 samples taken evenly from 10 demes. Both Nm (0.001, 0.002, 0.004, 0.008, 0.016, 0.032, 0.064, 0.128) and d/D (1, 0.5, 0.2, 0.1, 0.05) were varied and 100 replicates simulated for each parameter combination. The results agree with those of Papadopoulou et al. (2008) in general, in that the chance of detecting significant clustering under the MYC model declines with increasing migration rates. However, inspection of Figure 2a shows that the robustness of the MYC method depends significantly on the sampling scheme. With decreasing d/D, the chance of detecting significant clustering in the face of high migration rates increases drastically (Fig. 2a). The main effect of migration at the beginning of the coalescent process is then to randomly move lineages into unsampled demes, thereby creating additional clusters and increasing the support of the MYC model. For instance, if only every 20th deme is sampled and Nm = 0.064, the chance of detecting significant clustering is still >0.8 (Fig. 2a). The overall excess of clusters detected by the MYC method matches the theoretical prediction for the number of lineages escaping coalescence in their local deme (Fig. 2b) (Wakeley 1998, equation 32). As expected, the fit to the prediction (which neglects the chance of migration to a sampled deme during the scattering phase) increases with decreasing d/D. In the extreme case of complete sampling (d/D = 1), the number of inferred clusters is slightly lower than the number of demes (the gray dashed line in Fig. 2b) because escaping lineages necessarily land in sampled demes. The results agree both with intuition gained from the separation-of-timescales arguments as well as earlier simulations (Wakeley 1998) in that d/D does not have to be very small for strong clustering to emerge in the face of migration. a) The proportion of genealogies with significant (P < 0.05) clustering under the MYC model plotted against the scaled migration rate. Different colors correspond to different sampling schemes, that is, proportions of sampled demes, d/D: Orange = 1, red = 0.5, green = 0.2, blue = 0.1, and black = 0.05. In each case, genealogies were simulated for a total of 100 sequences. 10 samples were taken from each of 10 demes. Each point is based on 100 replicates. The orange line corresponds to the complete sampling scheme assumed by Papadopoulou et al. (2008). b) The average number of clusters inferred by the MYC method for different sampling schemes. The upper dotted line is the theoretical prediction for the number of lineages at the end of the scattering phase in the limit when d/D tends to zero. Given the large effect of the sampling scheme, how realistic is the assumption of incomplete sampling? First, geographic sampling is hardly ever complete in practice. This is true in particular for most barcoding data which are rarely collected with a particular sampling scheme in mind (but see Pons et al. 2006; Papadopoulou et al. 2008), and it has been argued before that the “barcoding gap” may in part result from incomplete spatial sampling (Moritz and Cicero 2004). Second, there are biological reasons why the kind of completeness required for the MYC method to be reliable may be impossible to achieve in practice. What governs the formation of clusters is not the population structure at the time of sampling but rather the sum of population structures that have affected the ancestral process of the sample in the past. The symmetric island model considered here is the simplest possible model of structure. In more realistic metapopulation models, demes are transient so that lineages may spend the majority of their history in demes that have subsequently gone extinct and can therefore not be sampled. Thus, increasing the geographic scale of sampling does not necessarily get around the problem. Considering that separation of timescales have been applied to a variety of models of structures (Wakeley 2004; Wilkins 2004; Matsen and Wakeley 2006), the main result is likely to hold in general. For instance, an analogous argument can be made for samples from a population in a continuous 2-dimensional habitat (Wilkins 2004). In this model, there is no discrete underlying structure at all so any observed clustering must be spurious. However, if a sample is taken from a set of random locations, one would expect a pattern similar to that observed in the island model. At the beginning, lineages either coalesce quickly in their neighborhood or escape by chance, in which case coalescence takes a much longer time on average. Again, the resulting clusters would only partly correspond to sampling locations with additional clusters being created by migration during the scattering phase (see Wilkins 2004, figure 4). In conclusion, the method of Pons et al. (2006) delimits essentially random clusters when applied to samples from a single island model population if d/D is low. Similar behavior is expected under any model of geographic structure as long as there is a considerable fraction of unsampled space and a separation-of-timescales exists. This is particularly worrisome considering the envisioned application of the MYC method to high-throughput mtDNA profiles (Pons et al. 2006). Such mass samples are likely to contain both individuals from truly isolated clades or species and structured populations connected by gene flow, making it even harder to distinguish between the two types of clusters. Taken together, the results cast serious doubts on the usefulness of mtDNA barcodes as a scaffold for an automated DNA taxonomy (Pons et al. 2006). The stochastic nature of both migration and lineage sorting requires multilocus data, exhaustive geographic sampling, and realistic models, which can deal with the expected incongruence between gene genealogies to delimit meaningful taxonomic units from sequence data (Edwards 2009). However, this remains a difficult task even for a very modest number of taxa (e.g., Knowles and Carstens 2007) and is incompatible with the notion of a DNA taxonomy based on a single locus. Given the ubiquity of population structure in nature, the number of potentially detectable clusters in mitochondrial barcode data is likely to vastly exceed that of meaningful taxonomic units. K.L. is funded by the Biotechnology and Biological Sciences Research Council. Many thanks to Nick Barton for much helpful advice and encouragement. Thanks to Jerome Kelleher for modifying MS; Marianne Elias, Brian O'Meara, and one anonymous reviewer for useful comments; and Anna Papadopoulou, Alfried Vogler, and Tim Barraclough for taking the time to discuss this problem.
No takes yet. Share an insight, caveat, or question.
Konrad Lohse (2009) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: