This editorial refers to ‘Cardiovascular disease burden from ambient air pollution in Europe reassessed using novel hazard ratio functions’†, by J. Lelieveld et al., on page 1590. During the severe smog event in London in the winter of 1952, which is now believed to have caused some 12 000 deaths, PM10 levels approached 3000 μg/m3, 60 times higher than the current EU air quality standard 24 h limit.1 , 2 The widespread concern among Londoners that developed during this event is believed to have catalysed the passage of clean air regulations soon thereafter. These and subsequent regulations have delivered huge improvements in air quality across Europe over the past 60 years, giving many people a sense that health threats associated with air pollution are a thing of the past. However, research increasingly suggests that air pollution, in particular PM2.5, continues to have an enormous impact on public health, in Europe and beyond. The work of Lelieveld et al.3 in this issue of the European Heart Journal presents a startling new picture of the impact of air pollution exposure on mortality in Europe, claiming that 790 000 excess deaths are attributable to air pollution each year. The excess mortality estimates from the new model employed by Lelieveld et al., called the Global Exposure Mortality Model (GEMM), are approximately double those produced by the model used by the World Health Organization.4 GEMM was developed previously by Burnett et al.,5 providing hazard ratio (HR) functions that can be used to estimate PM2.5-attributable disease-specific mortality. The GEMM-estimated HRs are informed by the relationship between PM2.5 and mortality in 41 epidemiological studies across 16 countries. GEMM improves on previous approaches to estimate such HRs because it includes studies from China where observed exposures are high, meaning that HRs at high exposure levels are informed by real data and do not have to be extrapolated or based on smoking data. This leads to larger HRs than those from older models, particularly at high PM2.5 concentrations.4 Lelieveld et al. use atmospheric chemistry modelling software called EMAC to estimate gridded annual average PM2.5 and O3 exposure levels globally (grid cells ∼122 km2). They then plug the PM2.5 estimates into the GEMM HRs and the O3 estimates into the HRs of Jerrett et al.6 to estimate global excess mortality rates attributable to these exposures (compared with a counterfactual scenario in which all PM2.5 exposures were <2.4 μg/m3; counterfactual O3 exposure used for comparison is unclear). They provide a detailed analysis of Europe in which they estimate excess mortality for all diseases combined and for five specific diseases: lower respiratory tract illness, chronic obstructive pulmonary disease, lung cancer, ischaemic heart disease, and cerebrovascular disease. Lelieveld et al. make important contributions to the literature with their novel combination of several state of the art tools for air pollution and health research, and with their deep dive into the health impacts of pollution in Europe. Thanks to recent advances in atmospheric chemistry modelling softwares such as EMAC, pollution exposure levels can be estimated globally at moderately high spatial resolution (EMAC can estimate for grid cells as small as ∼50 km2). With these pollution estimates and improving methods for HR estimation such as GEMM, we are now well positioned to estimate the global burden of disease due to pollution. Importantly, these tools also allow for assessment of the health impacts of pollution in areas of the world where pollution monitoring and epidemiological studies of pollution and health are scarce or non-existent. The use of the GEMM HRs by Lelieveld et al. bestows unique credibility on their results for regions with high pollution exposures,7 thanks to the representation of high exposures in the data used to construct the HRs. GEMM also improves on older models in its handling of non-linearities and uncertainties. GEMM employs raw data from 15 epidemiological studies to directly model the functional form of the relationship between PM2.5 and mortality, and captures uncertainties surrounding this functional form using ensembling approaches. Then, through meta-analysis of the results from these 15 data sets and 26 additional studies from which aggregate results are available, both within-cohort variability and across-study heterogeneity are captured in the HR uncertainties. While Lelieveld et al. have made creative use of several cutting-edge resources to provide new evidence of the devastating impact of pollution on health, their work also brings to light many remaining data science challenges that must be confronted in order to reliably estimate uncertainties and make causal, population-level conclusions about pollution-attributable mortality at the continental or global scale. GEMM applies advanced techniques to account for uncertainties, yet the diverse and abundant sources of epistemic uncertainty in this setting call for even more complex modelling approaches. We applaud the authors for acknowledging that their approach probably underestimates uncertainties and for discussing some of the neglected uncertainties, including uncertainty regarding confounding adjustments and generalizability of the HRs. To explain the sources of uncertainty at play in this analysis, we distinguish: measurement error (MEA), interpolation/extrapolation error (INT), sampling error (SAM), and modelling error (MOD). We consider the most notable sources of uncertainty introduced in each stage of the approach of Lelieveld et al. First, global/regional pollution exposures are estimated from a model, and uncertainty is likely to accompany the exposure estimates, due to MEA in data inputs to the atmospheric chemistry models, uncertainty about model parameters (MOD),8 and the aggregation in the gridded estimates (INT). In the next stage, the estimation of the exposure–response models and HR function, each epidemiological study utilized will have its own sources of MEA and SAM. Moreover, there is study-specific uncertainty about the structure of the exposure–response model (MOD), including uncertainty about the functional form of the exposure–response relationship, the set of confounders adjusted for, and the form in which confounders should enter the model (note that the strong assumption of no effect modification is needed to make later stages in this approach valid). Inter-study heterogeneity, a type of SAM in this context, also impacts the estimation of the HR function. In the final stage, global/regional pollution exposure estimates are plugged into the HR to estimate pollution-attributable mortality. Uncertainty arises in this stage through the generalization of the HR to the global/regional population, which is probably not fully represented by the study populations used to construct the HR. Generalizing results generates additional SAM and requires extrapolation if these populations experience exposures outside the range observed in the studies (INT). See Figure 1 for a summary. Sources of uncertainty introduced in each stage of the standard modelling approach for estimating pollution-attributable mortality at the global and regional scales. In the model used to estimate global/regional pollution exposures, uncertainty arises from measurement error in the input data, unknown model parameters, and the aggregation of the exposures to grid cells. The exposure–response models to estimate the hazard ratio (HR) function are constructed using many epidemiological studies. Here, uncertainty is introduced through measurement error in pollution and confounder data, within-study variability, unknown within-study model form, and inter-study heterogeneity. When global/regional pollution exposure estimates are plugged into the HR to estimate total pollution-attributable mortality, additional uncertainty is generated. This is because the data used to create the HR probably do not represent the global/regional population and because the HR function may need to be extrapolated to accommodate exposures outside the range observed in the epidemiological data. Moving forward, greater emphasis should be placed on the development of methods that rigorously quantify the uncertainties introduced at each stage and propagate them to subsequent stages, so that they are reflected in the final uncertainty estimates. Such methods would substantially improve reliability and inference. In generating pollution estimates, atmospheric chemistry modelling softwares do not currently have the capacity to produce reliable uncertainty estimates. However, recent applications of Bayesian and machine learning methods7 , 9 to estimate pollution concentrations on a large scale present a promising way forward for accurately estimating uncertainties in this stage. The use of more flexible Bayesian exposure–response models (where raw data are available) would reduce uncertainties surrounding model specification,10 , 11 and these could straightforwardly accommodate measurement error in pollution/confounder data.12 When the HRs and global/regional pollution exposures are combined to estimate pollution-attributable mortality, Bayesian sampling or Monte Carlo techniques could be applied to account for the uncertainties in both. While weighting approaches might help to account for uncertainty regarding generalization of the HRs to the global/regional population, generalizability is an active area of research,13 and more work is needed to develop generalizability concepts and methods at the global scale. The GEMM HRs are created using observational epidemiological data, thus the relationship between pollution and mortality in these data is subject to the threat of confounding. In observational data settings, strong assumptions regarding proper adjustment for confounding are needed to establish causality. In many if not all of the studies informing GEMM, confounders are adjusted for by including them as linear terms in a regression model. Inferring causality in this setting requires the assumption that all confounders are observed and that the relationship between each confounder and the outcome is truly linear. Additional assumptions must be made regarding the generalizability of causal effects to the global/European population in order to interpret the results of Lelieveld et al. as causal. Future work on this topic could develop or integrate statistical methods that relax some of the strong assumptions required to draw causal, population-level conclusions from the results of Lelieveld et al. Recent work has demonstrated that the application of causal inference tools in air pollution epidemiology can increase the plausibility of the assumptions required for causality.14 Moving forward, exposure–response models can be formulated within a potential outcomes framework, and novel methods for estimation of causal exposure–response functions with minimal modelling assumptions15 could be extended and applied to estimate HRs. In conclusion, while Lelieveld et al. have demonstrated that we are making great progress towards accurately quantifying global/regional pollution-attributable mortality, new statistical methods are needed to better estimate uncertainties and to relax the strong assumptions relied upon in the current approach. Recent developments in Bayesian causal inference methodology present a promising direction for achieving these goals. This work was supported by the National Institutes of Health (grant nos 5T32ES007142-35, R01GM111339, R35CA197449, R01ES026217, P50MD010428, DP2MD012722, R01ES028033, and R01MD012769); the Health Effects Institute (grant no. 4953-RFA14-3/16-4); and the Environmental Protection Agency (grant no. 83615601). Conflicts of interest: none declared.
No takes yet. Share an insight, caveat, or question.
Nethery et al. (2019) studied this question.