Vaginal microbicides are topical applications used to help women protect themselves from sexually transmitted infection. Consideration of their role as a public health measure against human immunodeficiency virus (HIV) in particular began in the late 1980s. 1 By then it had become clear to many in the HIV field that advising women to insist that their partners use condoms was seldom realistic; it assumes a power that few of the women most in need possess. The underlying aim of the vaginal application was to provide a woman with a means of protecting herself without requiring her male partner to participate or even to know. By February 2002, plans for at least five trials of candidate microbicides were approaching final form. At the time of this writing, none of these has entered phase 3, the final test in the field trial itself. 2 It seems to us that the design strategies of those studies closest to implementation are in some respects deficient. Thus, the moment is timely for raising epidemiologic concerns. In what follows we address two issues: one is the choice of controls, and the other has to do with problems and benefits of contemporaneous trials at multiple sites. Randomized controlled trials (RCTs) of a new pharmaceutical product are mounted only after preliminary work (including tests on laboratory animals) has supported their likely efficacy. Phase 1 is conducted in human subjects primarily to ensure safety and to guide dosage. Safety is further assessed in phase 2, and the appropriate trial procedures are devised and developed in the “at-risk” population itself. In phase 3, investigators can begin to evaluate the product with continued monitoring for safety and effectiveness in the field. In the United States, each step is monitored by the Food and Drug Administration (FDA). Control Groups Although several aspects of design for microbicide trials have received close attention, the critical choice of a control arm still engenders contention. Two distinct types of controls have been proposed. One is the standard of choice, a double-blind placebo arm in which neither the participants nor the researchers are permitted to know the treatment assignment. The placebo should be an inert substance indistinguishable from the experimental product in physical characteristics. In several of the new trials proposed, however, an additional control arm receives only condoms together with counseling on their use. In an editorial accompanying a recently published microbicide trial, 3 at least one experienced statistician approved a design in which a single control arm consisted of condoms without placebo. 4 We have encountered two grounds for the use of a “condom-only” arm: first, to ensure a determinate result; and second, to enable extrapolation to the population at large (the “real world”). In what follows, we voice our doubts about whether this tactic can be relied on to fulfill either of these intended purposes. Concern about the selection of appropriate controls first arose in 2000 in the wake of the UNAIDS COL-1492 trial results. 5 In an untoward result, HIV incidence appeared slightly higher among subjects assigned to the experimental gel (and condoms) than among those assigned to the placebo (and condoms). The interpretation of this result remains uncertain. To explain it, some researchers (including authors of a World Health Organization [WHO] report 6) have suggested that the experimental gel may not necessarily have caused harm. Instead, the placebo itself, not being the “inert” product it was assumed to be, could have had a preventive effect on HIV. To circumvent this design hazard, some measure to gauge whether the placebo was indeed inert would perhaps help. It is apparently to this end that the several impending American-sponsored trials (although so far not the British trial) have adopted the additional “condom-only” control arm. Thus, the proposed trial design incorporates dual controls: a placebo arm and a “condom-only” arm. In effectiveness trials, these are to be compared with one or more test products, the putative microbicides. In evaluating the effectiveness of an experimental product, a design that adds new arms has a substantial impact on the time, resources, fieldwork and quantity of data in studies that are already very large. The implications of such an approach in testing microbicide effect therefore require careful scrutiny, not least with respect to inference. Table 1 presents a set of the potential outcomes of testing a hypothetical experimental microbicide in a randomized trial of the kind proposed. Relative risks of HIV incidence in the microbicide arm are compared with those in the placebo and the “condom-only” arms. In rows A, B and C, the relation of the putative microbicides to both placebo and “condom-only” arms is the same. Interpretation is straightforward. The placebo arm provides the more credible reference for microbicide effect. At best, the “condom-only” arm supplies secondary confirmation.TABLE 1: Hypothetical Outcomes and Relative Risks (RR) of a Randomized Controlled Trial Comparing an Experimental Microbicidal Product with Placebo and to “Condom Only”* †In all the remaining rows of the (D to I), the relative effect of the putative microbicide differs depending on the control comparison. Each of these rows allows a discomfiting number of interpretations, some more and some less plausible, but no one of them secure. No great difficulty resides in interpreting the standard comparisons of putative microbicide vs placebo. But the third “condom-only” arm is a metaphoric joker. Addition of an unblinded group compromises the neutrality between arms with respect to behavior conferred by randomization. Of particular note are rows G and H, in which the putative microbicide appears either equivalent to or worse than placebo (although both are better than “condom only”). In the COL-1492 trial (with no “condom-only” comparison) the result was indeterminate: either the placebo was more effective than the microbicide, or the microbicide was harmful. It is this disconcerting result that fueled calls for the inclusion of a “condom-only” arm in future microbicide trials. We turn now to consider whether such a step would serve a determinative result. Because of the open nature of trial assignment in comparisons of a topical agent vs “condom only,” outcomes could very well reflect differences in behavior rather than interventions. Across an open trial, reports of behavior and especially reports of sensitive sexual activity could vary with trial assignment. Thus, retrospective reports of sexual behavior and of condom use could be of questionable validity. Variability in actual behavior, in reported behavior and in interaction between these could yield three uncontrolled sources of variation between the participant groups across the arms of an open trial. We see no way in which the potentially confounding effects of such sources of variability can be adequately controlled. In comparisons across properly blinded groups, such doubts should not arise. A second argument for the use of a “condom-only” arm assumes that the trial comparison involving such an arm will better represent the parent population untouched by intervention. Supposedly, this would measure the protective effect of microbicides on women's risk of HIV infection in the “real world.” Randomized controlled trials do not aspire to represent experiences and effectiveness in the “real world.” In the ideal, such trials test the effectiveness of a particular intervention by comparisons under arbitrary conditions created to conform to the imposed requirements of the research and differing only according to the intervention under test. Departures from et ceteris paribus (literally, all other things being equal) introduce the biases and confounding that randomization and blinding, the essential features of randomized control trials, are intended to prevent. In a trial, two things act selectively and virtually preclude true representation of a specified population at large. One is the ethical requirement to enroll consenting volunteers, and the other is variability in readiness to comply with differing forms of intervention. Moreover, the generalizability of any clinical trial result will be problematic for reasons aside from the nature of the control arm. Under the conditions of research and of the “real world,” measures of effect are bound to be different: first, behavioral changes are to be expected after the introduction of such new types of intervention as the microbicide; second, a gulf separates the rigor and demands of properly implemented clinical trials on the one hand and the delivery of health services on the other. Strictly, the goal of phase 3 microbicide trials is to test whether a microbicidal product lessens the chances of HIV transmission. We conclude that the inclusion of a “condom-only” arm will not contribute to this goal, and could detract from it. Should any experimental microbicide appear ineffective in comparison with the best available placebo, then the experimental product is unlikely to be effective for widespread HIV prevention. Given limited resources and the tremendous urgency of identifying an effective microbicide for HIV prevention, the additional cost and time associated with a “condom-only” control arm is difficult to justify. Multiple Trials The second issue concerns the multiple contemporaneous trials of microbicides that we anticipate will be entering the field in 2003. 2 This is a welcome development; the devastation of the HIV epidemic justifies pressing forward as quickly as possible to a judgment about each of the proposed test substances. Sponsored and funded by different institutions, the trials emanate from different countries and operate at different sites. Given this multiplication, the issue of concern is how best to achieve the common goal of a proven and safe microbicide. At the time of this writing, we anticipate there will be at least five trials, each testing different substances. We have no way at this stage to choose among the five or six substances to be tested. Several of the test microbicides are rather similar in supposed modes of action or differ only in dosage. To distinguish among them, however, demands the expensive and laborious task of conducting field trials. Each trial will require a large population to be tested. Prophylactic trials must assemble an unaffected population in which the numerator will be the cases not prevented. The large numbers required are a difficulty alleviated only by very high incidence rates or by highly effective prevention. There are few populations in which enough incident HIV cases can be expected to arise at any one site over a reasonably brief observation period. To achieve the required numbers, several of the individual trials therefore incorporate three or four different sites, each using similar designs. Given this necessity for several sites, statistical adjustments for differences across sites will of course be required. Yet adjustments are unlikely to assure complete control of factors that may yet produce differences between sites. Hence, in a trial with several sites the populations should be as homogeneous as possible. The same principle holds for assessing the performance of different substances. If their effects are to be compared across different trials, homogeneity among trials in all attainable ways is desirable. In practice, neither of these economies is achievable in full. The difficulties of execution, the variation in expected incidence rates, and the large administrative and funding issues preclude the ideal situation. This is not to say that variations in circumstance and procedures can yield no advantages. The issues are well stated in John Stuart Mill's “canons,” logical strategies he devised for inferring causation. 7,8 Two of the canons are relevant here. His second canon, the “method of difference,” can be briefly restated: the situations compared are alike in all variables but one. This is the rigorous standard to which randomized controlled trials aspire and for which we have argued above. Mill's first canon, here briefly restated, is the “method of agreement.” The situations compared are consistent in having only one circumstance in common. In other words, if a putative microbicide has the same effect across different situations, the result offers grounds for a causal attribution. This canon is a saving grace for causal inference in observational epidemiology. Given adequate rigor, even in different situations replication of a result adds strength to a causal attribution. Nevertheless, in randomized controlled trials, the more similar the testing procedures are to each other, the more it will be possible to consider the results of the different trials as reinforcing the results of each, whether confirmatory or rejecting. In this light, the choice of placebo becomes a key issue in comparisons across trials. Every effort should be made to employ an identical substance as placebo. For each trial, the placebo should look, smell, taste and be of the same viscosity as each putative microbicide. Admittedly, the degree to which the vehicles for a test product differ across trials could make it difficult to achieve complete uniformity among placebos. This does not relieve the trialists from making the effort to strive for uniformity. To this end, the composition of each placebo should be knowledge freely shared across trials. In practice, too, several of the current trials deliver the test substance and placebo in applicators, which confers some flexibility in the degree to which the appearance of each needs to be identical. Given the unplanned but welcome contemporaneous introduction in current trials of several putative microbicides, there is every reason to provide to the extent possible a standard that would permit direct comparisons of effect. These trials are critical in dealing with the worldwide contagion of HIV/AIDS. The need is pressing to demonstrate the effectiveness of interventions, for microbicides as for vaccines. Epidemiologists have a duty to produce sound study designs. Statisticians are an invaluable and indispensable resource in such endeavors. They can advise on and perhaps perform the most appropriate analytic techniques, as well as those statistical refinements (multiple adjustments, metaanalyses, etc.) that help to correct for bias and the like. Nevertheless, epidemiologists have to know that designs flawed in the first place will inevitably weaken inference and interpretation. Epidemiology lives or dies by the rigor, integrity, economy and good judgment applied to research design.
No takes yet. Share an insight, caveat, or question.
Stein et al. (2003) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: