This is the book we have been waiting for—occasionally a book is a classic by the time it is published and this is it. We have needed this for a long time; phylogenetics has grown enormously over the last two decades and it is almost impossible to keep a proper overview. By my estimate (not ML, I'm afraid) this book has around 1000 references. There is no way that ordinary mortals can keep up with such a literature on methodology alone, let alone work back through it to find sources of ideas, or possible alternative approaches that need developing. The breadth is very wide with all the main expected topics. Numbers of trees, parsimony algorithms, distance calculations, Markov models for sequence evolution, likelihood and Bayesian methods, bootstraps, likelihood ratio and other tests, consensus, and coalescents. Such a full treatment requires over 600 pages, which is not surprising given that many less common approaches are also covered including compatibility, invariants, Hadamard transforms, restriction site data, quantitative characters, tree shapes, and even to ways of drawing trees. A strength of the book is that it presents the theory behind the methods in a clear manner. A lot of thought has gone into presenting the equations in ways that quantitative biologists can gather the key points without being lost in additional mathematics and proofs. Okay, perhaps the others can skip over the equations and follow the text, but that would miss a lot of the benefit. The book has clearly been put together over a long time period, it is of mature vintage. There has been time to fine-tune the wording, to look over how the equations and diagrams are presented. The extra time has helped get things right, sentences reworked, an obscure reference found, a gentle joke included. It turns out that Joe is the founder of the “It-Doesn't-Matter-Very-Much” school—get the tree right and shut up—with a healthy scepticism towards the importance of a formal classification. Overall the book has that comfortable mature feeling, no, not quite like mulled wine, more a really good mellow single-malt whisky with just a bit of bite—the Laphroig of the phylogenetics market. One theme that becomes apparent is that many ideas have been discovered or reinvented several times. This is not really surprising, estimating evolutionary trees accurately is a major problem in mathematical inference with the relevant theory coming from mathematics (graph theory, combinatorics, and Markov chains), statistics (likelihood, resampling, Bayesian), operations research (optimization, search heuristics), and computer science (complexity theory). No one (including mathematicians!) are expert in all those fields. Consequently anyone with a new idea to try out will especially gain from the book; there is just so much in it that will help get a quick start on how similar ideas have worked. Somehow a reviewer feels duty bound to find some reference missed, a factlet wrong, or similar minor misdemeanour. I definitely would have included Hartigan (1973) as a brilliant early piece of work; he proved that the Fitch algorithm for the length of a specified tree would always give the minimum length. On the surface it looked as if that calculation had to increase exponentially with the number of taxa—just look at the number of ways you could label the internal nodes of Figure 1A as the number of taxa grew. Then find your way though all possible paths to get the shortest. Surely, the calculation had to be exponential. To prove an algorithm that increased linearly with the number of species was amazing. So having found one reference missed I could feel pleased with myself, relax, and enjoy the book. However, I was soon deflated by a (favorable) reference to my own work that I had completely forgotten about! The moral of the story is not to try and catch Joe out, just enjoy the amazing resource he has assembled. (You still should check the list of corrections at http://evolution.gs.washington.edu/book/typos.html.) Perhaps the only change in emphasis I would like is to see more attention given to properties of the data—and therefore when method X might therefore be preferred over method Y. For example, there is still a strong tendency to say that maximum likelihood on sequences is ‘statistical’ whereas parsimony on morphological data is not. This seems wrong; both examples are ML estimators for different classes of data (see Fig. 1). An important difference is how we handle missing data for sequences and for morphology. We have information about the leaves (terminal nodes) of the tree, but not about the changes along the branches or internal nodes (see Fig. 1A). The brilliant thing about molecular data is that in treating it as a Markov model we do not have to estimate the character-state at all points along the branch. Integrating the rate matrix does this for us (Fig. 1B); all we need do is sum likelihood values over all character states at the internal nodes, giving maximum average likelihood (Steel and Penny, 2000). Averaging over missing data for (A) morphological data and (B) for molecular data. For molecular data (B) maximum likelihood need only average over all codes at the internal points, the Markov model handles the ways the data may have changed along each branch, this is MLaverage of Steel and Penny (2000). For morphological data (A), on current implementations without an easy justification for a Markov model, maximum likelihood has to consider what the data might have been at all points along each branch, and gives MLevolutioanry path, which is parsimony. Will still could do better, but the maximum likelihood estimator depends on the properties of the data. Averaging over missing data for (A) morphological data and (B) for molecular data. For molecular data (B) maximum likelihood need only average over all codes at the internal points, the Markov model handles the ways the data may have changed along each branch, this is MLaverage of Steel and Penny (2000). For morphological data (A), on current implementations without an easy justification for a Markov model, maximum likelihood has to consider what the data might have been at all points along each branch, and gives MLevolutioanry path, which is parsimony. Will still could do better, but the maximum likelihood estimator depends on the properties of the data. With morphological data a Markov model does not seem relevant. It's the data that are different, not the maths. We still have further to go here; Paul Lewis (2001) has a brave attempt at a new ML method for morphological data. Both approaches—parsimony and likelihood—are ‘statistical’; it is the data that's different. Myself, I'm a fan for using molecular data in phylogeny; then using morphological data to trace the real biological questions—morphological, ecological and physiological changes and adaptations during evolution. Speaking of properties of the data, we really also needed more on the importance of among site rate variation, and the consequences of not considering it. The properties of the data are as important as how the methods actually work. Okay, stop whinging. What of the future; we now have an outstanding summary of where we have been over the last 25 years. Perhaps it is foolish to try and be ‘wise before the event,’ but that is what fools are for. Take just one example, for the last 20 years phylogenetics has been dominated by sampling error—we have bootstrapping, convergence tests, Templeton, Shimodaira–Hasegawa tests, and on and on. In the past the universal problem seemed to be short sequences—not enough data. Suddenly we are getting much longer sequences (for example, Rokas et al., 2003), so why do we keep getting some particular trees wrong? We increasingly understand that, according to the current models we love and trust, sequences must run out of phylogenetic information after a few hundred million years (see Mossel, 2003). There is a comfort zone where sequences work well. Graduate students please their advisors by running simulations in this zone, and cunningly leave the time-scale as the proportion of mutations per site (not as the product of mutation rate and time, that would upset our belief in the infallibility of sequences). Our models are still incomplete by a long way, and there can be strong systematic biases in many data sets. Bootstrapping or Bayesian support values are no help if we are converging to the wrong tree. Is Bayesian support the art of being wrong—with confidence? How well does the best model fit the data, how much of the variability is explained by the model? Yes, there is a lot to do. But if Joe had stopped to write a chapter on possible future directions we would have had to wait much longer for this book—perhaps forever. The rest of the field keeps expanding and would then require more writing; we needed this book now! The advertisements suggest that Inferring Phylogenies could be used as a textbook. I doubt it—that's really publisher mis-talk. The book is a resource. It is absolutely essential for any laboratory. I am keeping mine at home so that I can always find it—keeps the fingers of grad students and postdocs off my copy. No, I am not being mean; we have bought two copies for the lab. It is hard to imagine how any lab could function without this book.
No takes yet. Share an insight, caveat, or question.
David Penny (2004) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: