Health effects are unknown for the vast majority of the >83,000 chemicals in commerce, creating challenges in balancing industrial needs against the complex landscape of health susceptibilities and exposures. The National Health and Nutritional Examination Survey (NHANES), a large-scale survey aimed at determining the prevalence and risk factors of major diseases, is increasingly used to postulate relationships between chemicals and adverse health effects in the U.S. population. The interpretation of these studies is complicated, however, by the ad hoc data mining approaches typically employed. Here we describe the use of frequent itemset mining for identifying exposure ⇒ health associations in data from the NHANES 2005–2006 cycle. From 9,440 dichotomized samples, 983 two-itemset rules were generated describing associations between markers of health and environmental exposure (lift >1, response threshold >97.5th quantile). A case study using parathyroid hormone levels to develop exposure-health effect hypotheses is presented. This case study demonstrates how association rules can be used in data mining to facilitate hypothesis development and improve traditional regression models by identification of potentially confounding variables even in the presence of missing information. Our approach is designed to enable more effective knowledge discovery of potential health impacts of environmental chemicals by facilitating comprehensive data mining and meta-analysis of the NHANES dataset. Long-term, our representation of the information allows for integration with other disparate data, such as known biological pathways, to address the current data gaps.
No takes yet. Share an insight, caveat, or question.
Bell et al. (2014) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: