Abstract The spatial resolution of environmental exposure and sociodemographic population data is often mismatched given limited publicly available population data that complies with privacy requirements for individuals. To address this limitation, we developed a novel matching algorithm to construct a synthetic population at the address‐level. To demonstrate how our approach can improve environmental justice (EJ) analyses and health impact assessments (HIAs), we examined sociodemographic patterns of residential proximity to major roadways in Greater Boston (Massachusetts) and HIA results, comparing our method with a random address allocation method. The synthetic population was developed at a census tract‐level using US Census microdata and combinatorial optimization methods and then downscaled to address‐level parcels by matching building attributes to synthetic households. We designated households within 50 m of a major road “high exposure” and households below state median household income “low income”.We found misclassification for individual households (21% of the high exposure/low‐income households in the matched data set were identified as such in the random allocation data set). We found modest aggregate differences in matched allocation (3.3% of low‐income households had high exposure) compared to random allocation (3.4%). In a HIA, the difference between random and matched allocation would be stronger when there is a strong interactive effect between a sociodemographic effect modifier and exposure on the outcome. Address‐level exposure assignment based on synthetic populations can provide more significant and nuanced health impact and EJ analyses. Our novel method can be applied to other regions of the US and expanded to other dimensions of population vulnerability.
Black‐Ingersoll et al. (Thu,) studied this question.