Rates of molecular evolution have been shown to vary significantly among nucleotide sites, loci, and taxa. In addition to these forms of rate heterogeneity, there is evidence that molecular rates vary with the timescale over which they are estimated. One of the most striking observations has been that of elevated mutation rates over very short timescales, such as those presented in studies of pedigrees (e.g., Howell et al. 2003, Millar et al. 2008) and mutation accumulation lines (e.g., Denver et al. 2000, Haag-Liautard et al. 2008). In contrast, much lower rates are observed over evolutionary timescales, as estimated in phylogenetic analyses calibrated with reference to paleontological or geological data. The disparity between rates of spontaneous mutation and evolutionary substitution can exceed an order of magnitude. Intermediate rates are expected between these two ends of the spectrum, but there has been disagreement over the exact form of the decline from the mutation rate to the substitution rate. Some authors have suggested that elevated mutation rates are very short-lived, perhaps persisting for only a small number of generations (Macaulay et al. 1997, Gibbons 1998). More recently, it was proposed that the estimated rate decays exponentially over tens to hundreds of thousands of years, producing a “time dependence” of rates, whereby the magnitude of the inferred rate depends on the age of the calibration used in the analysis (Ho et al. 2005, Ho et al. 2007c, Penny 2005, Ho and Larson 2006). Although some of the original evidence for this hypothesis has been challenged (Emerson 2007, Bandelt 2008), there has been a steady accumulation of empirical and theoretical support for a prolonged elevation of short-term rates (e.g., Genner et al. 2007, Burridge et al. 2008, Henn et al. 2009, Peterson and Masel 2009, Soares et al. 2009). This has included compelling evidence from analyses of ancient DNA (aDNA) in which the sampling times of the heterochronous sequences are able to provide calibrating information for estimating rates (e.g., Lambert et al. 2002, Barnes et al. 2007, Ho et al. 2007b, Hay et al. 2008, Subramanian et al. 2009a). In a recent critique, Debruyne and Poinar (2009) have claimed that the high rate estimates obtained in Bayesian analyses of aDNA data are an unintended consequence of analyzing short sequences. According to their “signal-dependent artifact” hypothesis, aDNA-based rate estimates depend almost entirely on the information content in the sequence alignment. The essence of the criticism is that the posterior distribution of the rate becomes so wide that the posterior mean becomes an upwardly biased estimator. This behavior has been noted in previous studies of heterochronous data with low information content (e.g., Ho et al. 2007c, Firth et al. 2010). However, Debruyne and Poinar go on to state that “the [rate] acceleration phenomenon is certainly of much lower magnitude than has been previously reported by Ho et al. (2005)” (p. 358). This is a misleading comparison because our study was based almost exclusively on analyses of modern DNA (isochronous sequences) using internal-node calibrations which, as argued by Debruyne and Poinar, are able to overcome the signal-dependent artifact. In fact, much of the evidence for time-dependent rates has come from analyses of isochronous data (e.g., Genner et al. 2007, Burridge et al. 2008, Henn et al. 2009, Soares et al. 2009, Papadopoulou et al. 2010). Resolving the concerns over Bayesian rate estimates from aDNA is important for several reasons. First, aDNA sequences typically range in age from 102 to 105 years, thereby filling a crucial calibration gap between the time periods covered by pedigrees (usually <102 years) and fossil-calibrated species phylogenies (usually >106 years). Second, because the ages of aDNA sequences can provide sufficient calibrating information for estimating rates (Drummond et al. 2003), these data make it possible to circumvent problems associated with choosing and implementing calibrations at internal nodes (Emerson 2007, Ho and Phillips 2009, Firth et al. 2010). In particular, terminal-node calibrations remove the need for assumptions about genetic divergence being correlated with the divergence of species or populations, which can be dubious because of the uncertainty posed by ancestral polymorphism (Charlesworth et al. 2005, Peterson and Masel 2009). Consequently, if rates can be accurately estimated from aDNA data, some insight can be gained into the underlying causes of time-dependent rates (Ho et al. 2007c). There is still uncertainty regarding the factors driving the time dependence of rates. Previous studies have considered the possibility of contributions from incomplete purifying selection, calibration error, sequencing error, aDNA damage, ancestral polymorphism, saturation, and model misspecification, among others (Ho et al. 2005, Ho et al. 2007c, Woodhams 2006, Henn et al. 2009, Loogväli et al. 2009, Peterson and Masel 2009, Soares et al. 2009, Subramanian et al. 2009a). Debruyne and Poinar have added to this list with their suggestion that mean posterior rate estimates are upwardly biased for data sets with low information content. Distinguishing among these various factors is crucial to future studies of recent divergence times, evolutionary rates, and the molecular evolutionary process in general. Below, we investigate the two major aspects of the critique by Debruyne and Poinar. The first of these is that the posterior mean provides a biased measure of the rate in Bayesian analyses of data sets with low information content. To examine this issue, we perform new analyses of sequence data simulated using known evolutionary parameters. We assess the relationship between sequence variation and estimated rate under a range of simulation conditions, including various rates and sequence lengths. The second major aspect of the critique by Debruyne and Poinar is that the artifactual rate estimates from aDNA data are governed by the information content of the alignments, as measured by the number of variable sites. Indeed, Debruyne and Poinar base their entire signal-dependence model on analyses of alignments of varying length, which they regard as a suitable proxy for information content. Although this might be appropriate for isochronous data, we argue that it does not provide the full picture for heterochronous data because the ages of the tips represent a crucial part of the phylogenetic and temporal signal. We propose that the amount of information contained in these ages depends on their structure and spread, including the length of the sampling interval in relation to the period spanned by the genealogy of the sequences (Drummond et al. 2003, Firth et al. 2010). To investigate this, we perform new analyses of 18 published aDNA alignments to assess whether the ages of the sequences in these data sets provide sufficient calibrating information for estimating rates. The results of these analyses show that most real data sets appear to have satisfactory temporal structure and signal. The results of our new analyses indicate that the “signal dependence” hypothesis has limited relevance to the majority of real aDNA data sets. Our results also suggest that the signal dependence cannot be regarded as an analogue to time dependence, unless one is willing to accept the validity of equating alignment length with temporal depth in aDNA data. Moreover, our results highlight the importance of other factors, including the distribution of sampling times, choice of population size prior, and the use of appropriate summary statistics in analyses of heterochronous sequence data that exhibit low variation. Here, we build upon a simulation study that was presented in one of our previous evaluations of Bayesian rate estimation using aDNA data (Ho et al. 2007b). Debruyne and Poinar have challenged the results of this study, criticising two aspects of our analyses. First, they argue that the rates estimated from the simulated data are more precise than those obtained from real aDNA data. Although this observation is correct, these results are an expected consequence of simulation-based analysis: the evolutionary models for nucleotide substitution and demographic history used in the analysis of the simulated data are chosen to match the conditions under which the data were generated. This is adopted as standard practice to make it easier to isolate the effects of the factor(s) of interest. The second criticism of the simulation study of Ho et al. (2007b) is that the substitution rate used in the simulations is too high, with Debruyne and Poinar stating that the rate is “25-fold the estimate of the substitution rate for the mt genome of vertebrates” (p. 350). However, this simulation rate was inspired by published estimates from the mitochondrial D-loop (Lambert et al. 2002, Shapiro et al. 2004), whereas Debruyne and Poinar compare this rate to that estimated from their elephantid data, which is based on whole mitochondrial genomes analyzed over a phylogenetic timeframe. Indeed, the vast majority of published aDNA data sets comprise sequences from the D-loop, which exhibits much higher mutation and substitution rates than does the rest of the mitochondrial genome in vertebrates. This also calls into question the design of the main analysis presented in their critique, in which subsamples from the complete mitochondrial genomes of woolly mammoths were taken to be representative of real aDNA data sets. Nevertheless, the high rate used in our simulation could be viewed as a legitimate problem if short-term rates were not actually elevated. This led Debruyne and Poinar to pose the question: “what would the accuracy and precision of the posterior rate of change be if a slower rate of substitution, in the range of the interspecific mitochondrial substitution rates (between 1 and 2×10 − 8 substitutions/site/year) were applied to simulate the same sequence data?” (p. 350). In response to this question, and to address some of their other concerns, we present the results of a detailed simulation study below. We conducted analyses of simulated aDNA data to investigate the performance of Bayesian rate estimation. The amount of rate estimation bias is quantified under various combinations of simulation rate and sequence length, including conditions that might match those commonly encountered in real aDNA research. We investigate the impact of varying the population-size prior, and we compare the performance of different posterior measures of the rate. Sequence evolution was simulated using Seq-Gen (Rambaut and Grassly 1997) on random trees generated according to a coalescent model with a constant population size of 105. Each simulated data set comprised 31 time-stamped, nonrecombining sequences, with ages of 0, 1000, 2000, …, 30,000 years. All sequences were generated according to the model of nucleotide substitution and with rate among and among were with different substitution rates − 8 − 8 and − substitutions/site/year) and sequence and the range of of aDNA data sets and conditions expected to sequence alignments with low information content. One data sets were generated for of sequence length and rate. from the substitution rate and sequence length, the simulations are to those in the sampling in our previous study (Ho et al. 2007b). rates were estimated from the simulated data sets using the Bayesian phylogenetic (Drummond and To match the simulation conditions, the substitution model was and a coalescent was chosen for the of was chosen for the substitution rate. of were obtained by with over a of with the first of as To compare different posterior measures of the substitution the and of the posterior rate distribution were for of were to for and sufficient sampling from the data the estimates of rate and population size are The population size can be in the estimation of rates, the data set is We this by sets of only in the population size population size to of population size a of and population size a of a range of that could be considered for vertebrates. that in these is actually as the of the population size and time in The performance of rate estimation among the sets of a of the of the population size the population size is to of estimates of rates are and The posterior interval of the substitution rate included the simulation at of the noted by Debruyne and Poinar, the mean posterior rate estimates that there is of the rate there is low information content or sequence in the data set substitution rate short sequence However, this bias in the more data sets. the posterior rate are the are biased than the The posterior which the a estimate of the to provide an measure combinations of substitution rate and sequence of results from the simulation study, the simulations with a population size of results were only from the that are in the of simulations in which the interval of the rate contained the of results from the simulation study, the simulations with a population size of results were only from the that are in the of simulations in which the interval of the rate contained the different the population size is an distribution of the analyses to posterior with not and with the population size and the rate The of analyses that to from to the simulation these are the appear to estimates of the substitution rate The interval of the than the population size was to included the simulation at of the estimates of the population size were obtained in the analyses that of However, in almost the simulation the rate was by the and This could be a consequence of the that analyses because those would have been the data sets with lower information content by a number of and producing lower rate this into it is to whether the estimation bias is or whether it results from a biased of the simulation the between mean posterior population mean posterior and for Bayesian analyses of data generated under different simulation conditions different rates and different sequence The results were obtained using an population size from to Each the results from analyzing from to by mean posterior population size The mean posterior rate estimate for the data set is also on the same a relationship with the estimated population Each simulation is a in the if the size for the posterior is which a of to the were from the posterior over a of with the first of as the population size is to a range of picture by the was with the simulation being from the interval to of the time The mean size of the interval is than in the analyses on the population the disparity as the number of variable in the alignment The posterior is the summary of the because the on population size also on the that can be taken by the substitution rate. In some the posterior distribution of the rate is to a the other the posterior mean to provide a estimate of the substitution rate it is possible that this is an unintended consequence of the population size the mean posterior rate might only be as a of the population size the substitution rate to in the of real information on rates in the data. This could some of the published rate estimates from aDNA sequence alignments, which have taken in of the low information content of the data. aDNA data sets vary in of their sequence and underlying substitution rates as as the temporal structure and of the would be to the information content in these data sets to whether they can estimates of substitution rates and divergence One of heterochronous data that is by the use of statistics et al. and in the analyses of information content by Debruyne and Poinar, is that the ages of the sequences form an important of the information content (e.g., Firth et al. 2010). This from the that the sequence ages are used for calibrating estimates of substitution rates. problem in analyses of heterochronous data is that rate estimates could be an of the sampling Here, we use a to investigate temporal structure in 18 published aDNA data sets. This data set the ages of the sequences and several previous studies of heterochronous data et al. 2009, et al. 2009, Subramanian et al. Firth et al. 2010). The analysis is able to provide some insight into whether the structure and of the sequence ages are sufficient to provide information on the rate underlying the evolution of the data the original rate estimate is in the data there is temporal structure in the original data set and the rate estimate cannot be et al. 2010). the Bayesian phylogenetic in (Drummond and we analyzed 18 published aDNA of the aDNA data sets analyzed by Ho et al. the alignment of woolly mammoths by Debruyne and Poinar, and a D-loop alignment et al. 2010). We data sets from the study by Ho et al. the and alignments contained too ancient sequences for the whereas the alignment is by the data set published by et al. The of the 18 data sets are in with in the original of aDNA alignments analyzed using the in the of sequence age of of aDNA alignments analyzed using the in the of sequence age of models were by comparison of Bayesian information with the number of taken as the size for the to the of the data models that a of were All data sets were as and a coalescent was for the and divergence All analyses were using a Bayesian demographic model et al. 2008). The demographic model size or Bayesian was chosen on the of of the In from the posterior were from a of with the first being as the number of was or in order to an size for the rate The sequence ages in of the 18 aDNA data sets were This was times for data set using the (Ho and 2010). Bayesian phylogenetic analyses were using the same as for the original data. data the demographic model was chosen to match that for the original data. The posterior rate estimates from the 18 data sets are shown in is to that among the data sets that the not rate estimates with wide In these the posterior rate was to the mean posterior rate not of substitution rates from a of aDNA data the first data the rate estimated from the original data set whereas the data represent the rates estimated from in which the ages of the tips were were to the if the mean posterior rate estimate from the original data set is not included in of the from the estimates from alignments that the estimates from alignments that the To investigate the of signal-dependent in these we considered the mean posterior rates in relation to the of the data sets from which they were estimated. Debruyne and Poinar that the mean posterior rate estimate be exponentially to the amount of information in the data as by the alignment We measures of information the number of sites, the number of variable sites, the number of sequences, and the of the number of and sequences in the alignment. the alignment of woolly which an and is of the D-loop alignment from the same we evidence that of these measures are to the mean posterior rate estimate in the aDNA data sets and in However, more than of the variation in rate estimates could be by an relationship with the age range of the sequences in data set insight into the temporal structure the data sets was gained the analyses. alignments the and In addition to the results presented in this study, previous analyses of aDNA from et al. and et al. have that these two data sets sufficient temporal information to estimates of substitution rates. the data sets that the the alignment is noted for low sequence with the observed variation by et al. The alignment is a small data set sequences over a short time et al. of the alignments and complete mitochondrial the Our analyses of simulated and real data show that the signal-dependent by Debruyne and Poinar is to have to the published rate estimates from aDNA data sets. Our simulated data sets a range of sequence and substitution rates, including those in real aDNA of the on population the posterior mean provides an estimate of the rate for the data sets simulated using a rate of − are to those of the mitochondrial D-loop in vertebrates. the alignments including those simulated using lower rates, there is some of estimation bias unless the population size is to simulation However, the rates used for the simulations in this study are low because they are based on phylogenetic short-term rates are actually as by the hypothesis of time-dependent rates, the estimation observed in this study might be to the majority of real aDNA The results of our simulation analyses that the posterior mean can be a biased measure of the substitution as by Debruyne and Poinar. However, the posterior estimates provide because the on the rates included the simulation The posterior which is to the a estimate of the to be the measure the data set has low information content. Nevertheless, it can a is for the population size for the substitution rate or age of the In of these it might be most appropriate to various of the posterior distribution of rates and other of interest. The measures the and to the same for the most data sets. to the by Debruyne and Poinar, our analyses suggest that sequence is not the the performance of rate estimation. factors, such as the ages of the sequences, and the structure of the underlying are also very important of aDNA data the population size is in some of the analyses a it is more to use an for the age of the than the population This is because population size and time are to estimate whereas the age of the can be inferred from or into the results of the the sampling ages of real aDNA data sets woolly D-loop, and woolly were to artifactual rate estimates using the This that some of the published aDNA alignments not sufficient temporal information to support estimation of rates and of sequence ages a for the validity of rate estimates from heterochronous data, including those from aDNA and et al. 2009, et al. 2009, Subramanian et al. Firth et al. 2010). is to that the alignment of complete mitochondrial genomes from woolly mammoths the This that the analyses by Debruyne and Poinar might be being based on a data set that is to posterior estimates information on the population size or There are several for the performance of the data First, the alignment only a small number of sequences. Second, the mitochondrial of mammoths has a with a very the two major et al. 2008). This is in the of the estimates that are obtained only the ages of the tips are used for calibration et al. 2008, et al. 2008, Debruyne and Poinar 2009). mitochondrial DNA has at an low a phenomenon that is in the genome 2008, et al. 2008). the substitution rate in has been much lower than that in which in have been more than other et al. This also calls into question the analyses by Debruyne and Poinar in which subsamples of the were to be representative of aDNA In short aDNA alignments have almost exclusively come from the D-loop, which is the most variable of the mitochondrial The used by Debruyne and Poinar to of the that are much more to data sets. Debruyne and Poinar that the bias to signal dependence can be overcome the of for at the of the this is possible appropriate in analyses of heterochronous data. rates were such analyses would need to be in a to the rate to vary between and et al. 2009). a molecular is as in the analyses by Debruyne and Poinar, rate different is as an a this would address the problem posed by time-dependent rates, it only does so by that the problem does not (Ho et al. 2007c). suggested to this problem is to the analysis to or sites, which are to a much of et al. 2009, Subramanian et al. et al. 2010). In internal-node calibrations are in analyses (Ho and Phillips 2009). Although we have that time-dependent rates are to be by a signal-dependent the obtained in the present study not published estimates of rates from aDNA data. estimates can be by a of other factors, including of the demographic model (Emerson 2007, Ho et al. 2007c, et al. 2009, and 2009, Subramanian et al. can in aDNA sequences, which can to biased estimates of rates (Ho et al. 2005, Ho et al. in several of the data sets have not been but their ages have been inferred by estimates from these data including the and be than those from data sets with However, higher rate estimates have been obtained from a wide range of aDNA data from a of with different demographic and that they not be with the high rates estimated in studies of pedigrees and mutation accumulation these results suggest that empirical and theoretical into the of time-dependent rates could be their very most aDNA data sets have low information content. Although the is as sequencing complete mitochondrial genomes to be from et al. 2008, et al. 2009, et al. 2009, Ho and short alignments are to a of aDNA studies in the In these the important question is not whether the information content is but whether it is sufficient for the analyses of interest. This was by the and the The authors and for and that the
No takes yet. Share an insight, caveat, or question.
Ho et al. (2011) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: