I would like to point out that there is a large body of literature, not referenced by Wickramaratne and Holford (1987, Biometrics 43, 751-765), concerned with the adequacy of control groups and measures of confounding. I am aware of three somewhat nonoverlapping branches of this literature. The first branch, exemplified by Rubin (1974, 1977), Rosenlbaum and Rubin (1983, 1985), Rosenbaum (1984, 1987), and Holland (1986), comprises the extensive work of these authors on problems of defining, detecting, preventing, and adjusting for inadequacies of the control group. Such inadequacies are formalized under the concept of nonignorability of treatment assignment (Rosenbaum, 1984). I suspect that Rubin's formalization is the best currently available for dealing with problems of causal inference, such as confounding. My only serious objection to this literature is that at some points it proposes to check for confounding by means of significance tests [see, for example, the test of strong ignorability in Rosenbaum (1984, ?5)]. In nonrandomized studies, it is the hypothesis that confounding exceeds a specified level (not the hypothesis of nonconfounding) that must be rejected before one can confidently proceed with inference regarding treatment effects. Thus, if one insists on doing a frequentist test for confounding, the logical choice is an equivalence test rather than a significance test (Greenland, 1989). Admittedly, an equivalence test requires one to parameterize the degree of confounding, but I view this complication as a benefit of equivalence testing: Given that some confounding is almost always present in nonrandomized studies, one should test whether the amount present is worth worrying about, rather than test a certainly false null hypothesis. A second branch, exemplified by Greenland and Robins (1986) and Robins and Morgenstern (1987), developed from the Miettinen and Cook (1981) article discussed by Wickramaratne and Holford. These authors examine the relation of formal concepts such as exchangeability (comparability) and collapsibility to intuitive notions of confounding. Like Wickramaratne and Holford, these authors point out the need to distinguish the phenomena of comparability and collapsibility when attempting to deal with confounding in a formal mathematical framework. Greenland and Robins (1986) also show that confounding and comparability can be defined without any reference to covariates, a point not mentioned by Wickramaratne and Holford, but which follows directly from Rubin's formalization. A third branch, exemplified by Gail (1986, 1988) and Chastang, Byar, and Piantadosi (1988), concerns adjustment for balanced covariates. In particular, Gail characterizes the class of models under which balance on a covariate (comparability with respect to the covariate) implies that the point estimator is over the covariate (here, collapsible means the unadjusted estimator is asymptotically unbiased for the parameter of interest). He also characterizes a subclass of this class in which failure to adjust for the covariate leads to invalid hypothesis tests and confidence intervals. The models in this
No takes yet. Share an insight, caveat, or question.
Greenland et al. (1989) studied this question.