As David Hunter puts it so aptly in his commentary,1 Pandora's box has been opened: genomics, transcriptomics, metabolomics, proteomics, interactomics, methylomics …. It seems like every day a new genome-wide technology is introduced. Is this to become “a diagnostic dream and an analytical nightmare?”2 What is the real value added that can be expected for molecular epidemiology? Much of the use of “-omics” technologies to date has been exploratory in nature.3 Analyses are typically performed in 2 stages using a training sample to search for patterns in a high-dimensional data followed by a validation sample to test the predictions of the model. A wide range of data-mining tools have been developed for this purpose, including hierarchical clustering, classification and regression trees, self-organizing maps, random forests, multifactor dimension reduction, and neural nets, to name a few (see Hoh and Ott4 for a partial review). I would like to explore a somewhat different paradigm: using these “-omics” technologies to inform hypothesis-directed pathway-based approaches to molecular epidemiology and to help direct genome-wide exploratory analyses into more promising directions. In either case, the basic idea entails using “-omics” data to inform the estimation of the parameters of a model for the epidemiologic data at hand.5 For example, in a candidate gene association study of a highly polymorphic gene like ATM, one might include in silico measures of evolutionary conservation and predicted effects on protein conformation or in vitro functional assays of the effects of each polymorphism6 as prior covariates in a hierarchical model for the relative risk of each variant.7,8 On a broader scale, one might use readily available genomic annotation data to prioritize the signals from an initial genome-wide association scan so as to improve the selection of markers to carry forward for testing in the later stages of a multistage study.9,10 If the multiple comparisons problem in a genome-wide association scan is not daunting enough, consider a study of the genetics of gene expression involving perhaps tens of thousands of expressed genes examined in relation to hundreds of thousands of single nucleotide polymorphism (SNPs),11 or all possible gene–gene interactions,12 or all possible haplotype associations13 in a genome-wide SNP scan. Here, the opportunity to exploit bioinformatics resources to focus the search has enormous potential that has so far not been tapped. In other circumstances, rather than relying on external bioinformatics databases, an investigator might want to apply some of these “-omics” tools directly to the samples from an epidemiologic case–control or cohort study to better characterize intermediate pathways.14 Here, one might think of the “-omics” data as providing the “missing link” among exposures, genes, and disease. The cost of these assays, difficulties in getting subject participation, or the need for special tissue preparations may preclude obtaining such data on all subjects in a large-scale study. Furthermore, the problem of reverse causation in a case–control design (the biomarker being affected by the disease or its treatment rather than the other way around) could make any simple case–control comparison of the biomarker meaningless. This suggests the need for some form of multistage sampling15 combined with modeling of the latent disease process.8 For example, one could use data from an appropriately designed substudy to build a model for the relationship of the biomarker(s) to their genetic and environmental determinants and apply the predictions of this model to the analysis of the full study in which biomarkers were not available.16 Alternatively, one could rely on the idea of “Mendelian randomization.”17 Here, to investigate the role of an intermediate variable on disease, one focuses instead on the associations of a gene with a biomarker for the intermediate variable and with disease separately. A variety of analysis approaches are available for such designs, but the challenges of extending them to genome-wide scale are daunting.18,19 There remain many challenges in the design and analysis of genome-wide studies using any of the technologies discussed here, and their real potential remains uncertain, even for the SNP association studies currently underway.20 Careful attention to the basic principles of epidemiologic study design, including avoidance of differential handling of case and control samples and confounding by ancestral origins, will be essential.21 Clearly, a fundamental requirement will be replication as has been well recognized in gene-association studies.22 Indeed, most gene expression and other genome-wide studies have generally adopted some form of training/validation 2-stage approach, although there may be circumstances in which a joint analysis is preferred.23 Despite these challenges, there is considerable enthusiasm, particularly in such fields as pharmacogenetics,24–26 in which the potential for personalized medicine in the form of genetically tailored preventive or therapeutic drugs is enormous. As exciting as these opportunities are, there will always remain an important role for classic risk-factor environmental epidemiology. Not every study need understand the molecular mechanism of an exposure–disease relationship for it to provide information that can be useful for informing important public health policy. As David Hunter points out, “the lure of new technologies should not distract us from following social, economic, demographic, or ecologic explanations for disease etiology.” ABOUT THE AUTHOR DUNCAN C. THOMAS is Professor of Biostatistics and Verna Richter Chair in Cancer Research at the Keck School of Medicine, University of Southern California, and Co-Director of the Division of Biostatistics. His research has focused on the development of statistical methods in epidemiology, particularly cancer epidemiology, occupational and environmental health, and genetic epidemiology.
No takes yet. Share an insight, caveat, or question.
Duncan C. Thomas (2006) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: