Key points are not available for this paper at this time.
Rothman and colleagues were invited to submit their piece to our recently established ‘Education Corner’, but on reading it we felt it merited discussion and debate.1 Those invited to comment considered that Rothman et al. had overstated their position but were basically correct.2–4 We are concerned that this notion will become accepted wisdom in epidemiology without its implications having been thought through, and feel that representativeness should neither be avoided nor uncritically embraced, but adopted (or not) according to the particular questions that are being addressed. The purpose of epidemiology is not simply to assess causal hypotheses.5 Rothman et al. elevate causal hypothesis testing to ‘a science’ and denigrate descriptive epidemiology as an applied practice and therefore ‘not science’. It is salutary to reflect that the main reason millions of dollars are poured into modelling the global burden of disease6 (GBD) is because epidemiologists neglected their responsibility to collect data on prevalence and incidence of disease from defined populations. Whatever one’s views are of the GBD initiative, it is certainly the case that evidence-based healthcare planning and policy require population-representative data. In causal studies, the fundamental concern is that lack of representativeness will introduce bias. Richiardi et al. in their commentary suggest that non-representative populations produce only weak bias in exposure-disease associations.2 Their conclusion is based on two of their own studies, one empirical and the other theoretical. Their empirical assessment compares associations of educational attainment and parity with two outcomes: low birthweight and caesarean section. They claim that non-representative internet sampling ‘does not necessarily introduce selection bias’. However, reviewing the data (Pizzi et al.'s Table 47) indicates that the non-representative sampling resulted in odds ratios that were lower for most comparisons, and the confidence intervals (CIs) of the relative difference in odds ratios between non-representative and representative sampling were consistent with effect sizes that would excite interest in many epidemiologists and result in a media frenzy. For example: ‘Any red meat you eat contributes to the risk (of death)’, claimed the lead author of a paper from the Health Professions and Nurses Health studies based on a relative risk estimate of 1.20, well within the limits of what was observed in these empirical estimates of the possible bias.8 Similar excitement was generated by a study that demonstrated a relative risk of 1.06 (95% CI 1.02–1.10) for cancer mortality in relation to a very large difference in intake of processed meat per day.9 The important point is that the distortions may be in any direction (which is unpredictable), and evidence from one empirical example does not necessarily apply to other causal questions. Richiardi et al.’s theoretical example10 uses Monte Carlo simulations to study the effects of selection into a study, claiming that the bias is small, particularly for relative risks between 0.5 and 2.0. However, they only examined scenarios where a single unmeasured determinant of the outcome also influenced the selection process, as they believe that ‘it is unlikely that multiple and independent important disease risk factors would affect the sample selection’.10 This is a surprising belief for these authors to hold, as their own empirical findings (admittedly published after this paper) show very clearly that multiple risk factors are indeed associated with participation (Table 1 in7). Why might the issue of multiple factors being associated with participation be important? Take the example of vitamin C levels and coronary heart disease (CHD) events: a large body of observational data from studies conducted at different times and places has produced a precise and repeatable estimate of an apparent benefit from higher vitamin C levels.11 But exploration of the complex confounding of the association demonstrates how multiple factors, operating across the life course, can lead to confounding strong enough to negate the apparent benefit.12–14 A large randomized controlled trial of vitamin C supplementation, in which such confounding should not arise, showed no strong evidence of any reduction in CHD events.15 A similar scenario, with conflicting results from observational and experiments, has been seen with respect to other antioxidant vitamins.16 The simple fact is that multiple factors do come together to generate sometimes sizeable non-causal associations, even after attempted statistical adjustment for confounding factors. The large American Cancer Society volunteer cohort is exactly the sort of non-representative study group that Rothman and others would consider fit for purpose—it is easy to follow up, participants are motivated to stay in the study and events are likely to be easy to count. In this volunteer cohort, high alcohol consumption was associated with a reduced risk of stroke,17 a surprising finding since the outcome included haemorrhagic stroke (for which alcohol might be expected to increase risk) and alcohol increases blood pressure which is a major causal factor for stroke.18,19 What type of heavy drinker would volunteer for a study about the health effects of their lifestyle? They are unlikely to be representative of all heavy drinkers in the population (e.g. they may be non-smoking, vigorous exercising, moderately wealthy epidemiologists) and the factors that make them non-representative will tend to render them at lower risk of stroke. By contrast, volunteers drinking moderately or less and non-drinkers may be more representative of moderate, low and non-drinkers in the general population. This non-representative cohort generates a potentially spurious result because many factors that are associated with the outcome of interest are also likely to be linked to self-selection into a study. In support of the argument for non-representative study groups, Rothman et al. state that scientific generalization is incongruous with representative sampling, only modestly reworking Rothman’s earlier views on the topic of representativeness.20 Using as examples animal experiments (where no attempt is made to sample from a population of animals) and randomized trials (where internal validity may be achieved by limiting recruitment to a narrowly defined group), an argument is developed that representativeness is counter-productive. The validity of animal experiments of pharmacotherapies has been widely questioned, as it has become clearer that, despite careful control for confounding factors, many of them get the wrong answers or are conducted with no intention of application in humans.21 Randomized trials of interventions that are primarily for use in older adults with multiple morbidities have been criticized for not recruiting participants more representative of those who will be treated in the real world.22,23 Trialists simply do not know (and certainly cannot control for them or generalize from restricted study groups) the complex confounding between age-related processes, co-existing disease and therapies, and the effects of a new intervention. Epidemiologists are not often capable of producing ‘general statements on nature’ and unfortunately more often report on associations in ways that imply that causal inference is being drawn.24 Rothman et al. consider that the best direction for epidemiology is to set up more studies that ‘control skillfully for confounding variables and thereby advance our understanding of causal mechanisms’. The UK Biobank study of 500 000 people is an outstanding example of a study (motivated initially by the desire to conduct large-scale genetic investigations) which implemented a demanding protocol in terms of measurements on participants where non-representativeness was inevitable. It is claimed for UK Biobank that ‘generalisable associations of exposure with disease can be obtained without including representative samples of particular populations’.25 The overall response rate of 5.5%26 is not prominently displayed on the UK Biobank website presumably because it is deemed irrelevant. From a purely gene-outcome association point of view, the study will, for the most part, be capable of yielding unbiased estimates of association, as genetic variants are unlikely to be associated with self-selection into the study and are not generally associated with confounding factors.27 The large sample size, relatively cheaply recruited, is a major advantage here. Non-genetic associations will have to be interpreted cautiously, as many variables of interest will be associated with participation and essentially volunteer samples may suffer from greater degrees of confounding than less selected samples. Once this data set (and all the others) is turned over to open access, it is inevitable that large numbers of environmental variable-outcome associations of small effect size—but very precisely estimated (see the confidence intervals on the relative risk of 1.06 for processed meat cited earlier)—will be published. It is this context that Rothman et al.’s hope for skillful control for confounding variables reflects optimism of a Panglossian scale: in many situations the degree of unavoidable measurement imprecision and inevitable unmeasured confounders renders reliable control unattainable. Rothman’s and colleagues’ stance bears some affinity with the justifiable excitement about the big data era we are entering.28 Proponents of this brave new world denigrate representative sampling in a way Rothman and colleagues would presumably applaud: ‘Reaching for a random sample in the age of big data is like clutching at a horse whip in the era of the motor car’.28 However, the promise of big data is explicitly not to identify causes; indeed, the ‘Ideal of identifying causal mechanisms is a self-congratulatory illusion’.28 We are in the realm of prediction, an example being the use of hundreds of variables from amount of TV people watch and the websites they visit to predict insurance risk. It’s much cheaper than the lab tests insurance companies often use, and does just as well, an exercise in ‘turning data into dollars’.29 But epidemiologists are surely not yet ready to abandon the difficult business of characterizing causality. Does any of this matter? Science is meant to be self-correcting; misleading findings will be exposed by further studies. Unfortunately, most of the studies that are capable of producing such findings share similar confounding structures (many components of which are not measured) and are only capable of making more precise but essentially meaningless estimates. Epidemiology that searches for causes of smaller and smaller effect sizes may become increasingly irrelevant when there is high profile contradictory evidence (including that pitting purely observational vs contradictory randomized controlled trial or genetic evidence). Growing awareness among the public, and research funders, of the impossibility of the robust detection and unlikely impact on public health of such findings might lead to less support for their generation. Moreover, the importance of a stochastic element to disease risk that is not epidemiologically tractable at the individual level is now apparent and argues for a re-orientation of epidemiology away from attempts to improve prediction of individual risk or search for non-existent additional causes, and towards making good use of genetic variation which may tell us about population-level modifiable causes of common diseases.30 In our first editorial for IJE in 200131 we quoted Reuel Stallones who in 1980 had memorably detected a ‘Continuing concern for methods, and especially the dissection of risk assessment, that would do credit to a Talmudic scholar and that threatens at times to bury all that is good and beautiful in epidemiology under an avalanche of mathematical trivia and neologisms’.32 In the pre-modern epidemiology world the focus was often on triangulating evidence from across as many sources as possible. Such evidence comes from a variety of sources, of which some are deliberately non-representative (for example the considerable value of twin studies for identifying potentially causal epigenetic influences on disease,33 or the follow-up of natural experiments) and some of which (including essentially 100% coverage population linkage studies) will be representative. Among these sources, large volunteer studies such as UK Biobank will be powerful tools, but will need to be combined with other approaches that allow strengthening of causal inference in observational data. We feel that representativeness should neither be avoided nor uncritically universally adopted, but its value evaluated in each particular setting. Conflict of interest: None declared.
Ebrahim et al. (2013) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: