The book is aimed at psychologists, not statisticians. The style is wordy and repetitious. It uses statistical words differently from us: ‘ANOVA’ seems to mean a hypothesis test using a single ratio of mean squares; ‘design’ a structured table of data. Each chapter concludes with exercises, some very simple. Lengthy appendices include end notes, references and answers to some exercises. The book says almost nothing about designing experiments. Most of the examples are not experiments because they are neither randomized nor controlled. Intrinsic factors are barely distinguished from allocated ones. There is almost no guidance on randomization apart from two good pages in Appendix C. The emphasis is on model comparisons: given one model contained in another, how much better does the latter fit the data, relative to its extra complexity? This seems to me to be a good, unsurprising, approach, but most psychology text-books, apparently, do not take it. However, the book does not really fit models, because it allows conclusions like ‘A is different from B but C may not be different from either’. Almost all of the data sets are made up, and chosen to make the arithmetic artificially easy. Part I introduces the conceptual bases for the rest of the book, and part II covers between-subject designs. Chapter 3 starts the ‘model comparison’ approach, obtaining the ratio between two mean squares as a reasonable measure of comparing two nested models in terms of ‘adequacy but simplicity’. The denominator is the mean square for residual from the larger, or ‘full’, model; the numerator is the mean square for the excess of this over the smaller, or ‘restricted’, model. So far, so good. However, this is not applied consistently when the treatment factor is quantitative (Chapter 6), and the only way of doing this for factorials (Chapters 7–8) is to neglect marginality. In spite of a heavy emphasis on significance tests, the book does recommend reporting the size(s) of effects as well! However, it never presents a sim-ple table of means with standard errors of differences, instead using several more complicated measures such as (μ^1−μ^2)/σ^. This would be meaningful if the only change from one experiment to another were a linear transformation in the scale of measurement; it is not sensible if different degrees of care in experimentation, or choice of subjects, lead to different sizes of experimental error. Some of these strange ‘measures of effect size’ are present because the response variable is only a proxy for what the experimenter is interested in, and someone else may use another proxy. In other words, the absolute difference in means may have no intrinsic meaning. I wonder whether any of the sophisticated models, or multitudinous tests, that are presented in the rest of the book are appropriate for such data. There are various conceptual errors in part II, concerning polynomials, orthogonality, random effects and the benefits of blocking, among others. Parts I and II of this book are like medium quality science journalism. The authors do their best to explain the topic, but they do not really understand it themselves so they persistently resort to citing other authors rather than present a clear, logical self-contained argument. They do not seem to know whether they are writing a text-book or reporting current research. I do not recommend reading it, but doing so may not do too much harm. In contrast, students who are exposed to parts III and IV will probably design their experiments badly and analyse their data inappropriately. Part III covers within-subjects designs. Because within-subject responses may be correlated in time, the book presents convoluted tests based on the assumption that the correlation of responses within subjects depends on the ‘within-subjects’ factor, whether this is time or an allocated treatment factor. Even without this complication, inconsistent procedures are given for allocated treatments. If there is a single treatment factor then time is taken into account, a row–column design is used and the model has additive effects of rows, columns and treatments. If treatments are factorial, then time is ignored, no advice is given on randomizing treatment order for each subject and the residual for each factorial effect is its own ‘interaction’ with subjects. Part IV tries to explain how to use restricted maximum likelihood for mixed models. The authors are long since out of their depth.
No takes yet. Share an insight, caveat, or question.
R. A. Bailey (2005) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: