Key points are not available for this paper at this time.
This second edition of this very successful book is a welcome update which should benefit both the rapidly growing user community and researchers who want to keep track of active recent developments and some hot topics in the growing area of statistical learning and data mining. When the book first appeared in 2002, it was a shock to see many techniques inspired from the computer science literature being embraced side by side with modern multivariate regression techniques, many of which were developed either by the third author, either individually or in collaboration with others, or the first two authors together. It was probably due to the late Leo Breiman who brought V. Vapnik’s work on the support vector machine to the attention of the statistical community, which is covered in Chapter 12 (‘Support vector machines and flexible discriminants’) (see Vapnik (1999)), which led to the revival of a key statistical area with a new catchy name, and inspired much recent interest in sparse modelling, though his work is very close to the familiar smoothing spline formulation by Wahba and others (see Wahba (1990)). The influence of computer and information sciences is apparent throughout the book. For example, the 1993 National Institute of Standards and Technology database as a benchmark inspired development of many methods discussed in Chapter 11 and Chapter 13. Vapnik (1999), page 173, discussed the superior performance of the support vector machine on these benchmark data, building on earlier work using variants of neural network methods. Many algorithms that are covered in the book are taken directly from the computer science literature such as those in Chapter 10 (‘Boosting and additive trees’) and many in Chapter 14 (‘Unsupervised learning’). Of the four new chapters, Chapter 15 (‘Random forests’) is based directly on Leo Breiman’s application of the ideas of bagging and boosting (which are discussed earlier in Chapter 8 and Chapter 10) to improve the tree methodology (which is discussed in Chapter 9). Chapter 16 (‘Ensemble learning’) covers more recent developments of applying boosting ideas to many other machine learning methods. Although choosing new topics to add to an already successful book is a difficult task, however, I feel that the choice of topics in Chapter 17 could have been done better, as it does not fit well with the rest of the book and does a poor service to the large area of graphical modelling. Instead, some mention of causality concepts such as Pearl (2009) may be more pertinent to the prediction flavour of the book. The last new chapter (Chapter 18, ‘High-dimensional problems: P≫N’) provides a welcome survey of some recent statistical thinking on selected topics on the theory of analysing high dimensional data. However, because this is an area of much current statistical research interest I think it left much more to be desired, which I shall explain later. Various other changes have been made throughout the book. However, the authors have chosen to leave most of the earlier chapters intact to preserve the familiarity for current readers. Indeed, Chapters 1–8 present a nice set of materials for a basic statistical introduction to regression modelling and model inference. The authors have clearly demonstrated their expertise and great insights in various places such as sections 3.4 (‘Shrinkage methods and Lasso’), 3.5 (‘Partial least squares’) and 3.8 (‘New developments on Lasso’) and Chapter 7 (‘Model assessment and selection’). Chapter 7 and Chapter 8 (‘Model inference and averaging’) comprise an important part of the book on the statistical principles for data mining, though model validation is rarely based on within-data assessment alone. Although I think that the strength of the book is to provide several statistical themes as guiding principles for algorithmic developments in statistical learning, further theoretical study of why or when these techniques should work in practice on high dimensional data sets is still very much in demand. Ripley (1996) contains more theoretical discussions which should complement this book well, in addition to Vapnik’s (1999) far-reaching discussions. Ye (1998) is an important reference which I am surprised is not cited in Chapter 7. Another example is the notion of the curse of dimensionality which is central to motivating many of the multivariate regression techniques that are discussed in the book, but which is poorly explained in pages 22–23 of the book. The recent book by Clarke et al. (2009) gives more attention to this issue. Because multivariate data often lie in a low dimensional manifold, as alluded to in section 13.3.3 (pages 471–475), the hopeless situation of the curse of dimensionality as discussed in pages 22–23 of the book may not happen in many commonly encountered situations. Indeed, real problems are more promising, and I hope that statisticians realize that it may be more realistic to assume much more structure with multivariate data. An example is functional data, a framework in which the shape invariance considerations in section 13.1.3 should fit very well. When there is strong underlying dependence between covariates in a multivariate regression model so that the underlying probability measure for the covariates may not have a joint density function and may have a much lower (possibly fractal) intrinsic dimensionality than the number of covariates, Lu (1999) and Ferraty and Vieu (2000) demonstrate a much more auspicious statistical theory of high dimensional non-parametric regression. I am not a fan of big books and the second edition has become heavy on the hand with over 750 pages. I hope that the authors will seriously consider eliminating some repetitive material and trimming down a little in a future edition (such as removing most of Chapter 8 and Chapter 17). Despite my misgivings, I think that the new edition should continue to enjoy the success of its path breaking predecessor, as more people are attracted to learning about this important area in this information rich age.
Zigang Lu (Thu,) studied this question.