First, Hanushek seems to question the validity of meta analysis. We, and many other scientists and statisticians, disagree. For example, a recent report of a committee of the Mathematical Sciences Board of the National Research Council (1992), studying problems of combining information in the empirical sciences, concluded that quantitative synthesis—meta-analysis—has gained increasing use in recent years and rightly so. Meta-analysis offers a powerful set of tools for extracting information from a body of related research (p. 2). Second, whether one admits it or not, taking evidence from studies (even if only a significance decision, a p value, or the sign of a coefficient) and drawing conclusions about the likely values of the parameters that the coefficients estimate is an inferential procedure. Inference procedures must have performance properties that are reasonably well understood if their results are to be credible. The fact that an inference procedure (such as that used by Hanushek) is vague does not usually exempt it from scrutiny, quite the contrary. Hanushek's statement that strong and consistent evidence meant to him that the preponderance of coefficients would have the same sign and be statistically significant, is a specification of his inference procedure. We admit that this inference procedure (vote counting) sounds sensible. However statisticians evaluate statistical procedures by deriving their properties. Hedges and Olkin (1980) showed that one would rarely obtain what Hanushek calls strong and consistent evidence even if the true values of the coefficients (the parameters) were positive and exactly the same in every study. Contrary to Hanushek's claim, Hedges and Olkin's results apply to a wide range of significance tests including t tests used for production function coefficients. Hanushek misunderstands the inference problem in re search synthesis in a subtle but important way. His inter pretation is that we want to know whether the estimated relationship in any study has the expected sign and is statistically significant. Statistical inference is used to draw conclusions about the true values (parameters) that characterize relations, not the (random) pattern of themselves. The failure to distinguish the pattern of observed results (es timates) from the parameter structure that generates them is a mistake that has led to a great deal of confusion in synthesis (e.g., to the appeal of vote-counting), and Hanushek seems to make this mistake. Third, Hanushek attributes to us the statement that none of our samples of estimates [is] appropriate for the statistical methodology because none are completely independent. We did not say this and do not believe it. Hanushek is aware of dependence in his data that could compromise his inferences, yet he apparently did not consider its consequences for the validity of his own conclusions. We did— by constructing a subsample of for each resource variable that was independent. Hanushek may wish that, unlike him, we had used full multivariate procedures in our meta-analysis to deal more efficiently with dependence (see Becker, 1992; Becker & Scpram, 1994). However, they require even better reporting of data than the methods we used (only six of Hanushek's publications provided enough information to use these methods; see Becker & Kamata, 1994).
No takes yet. Share an insight, caveat, or question.
Hedges et al. (1994) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: