A machine learning model using address-level discrete social determinants of health predicted diabetes prevalence better than the area deprivation index (testing adjusted R2 0.948 vs 0.381).
Cross-Sectional (n=2,369,365)
Machine learning models utilizing discrete address-level SDOH features significantly outperform traditional area deprivation indices in predicting tract-level diabetes prevalence.
Effect estimate: Adjusted R2 0.948
Introduction and Objective: Area indices used to assess SDOH’s impact on diabetes outcomes collapse many SDOH dimensions and flatten address-level variation, complicating study of SDOH pathways. We use machine learning on discrete address-level SDOH measures to identify pathways for diabetes prevalence. Methods: We used GPS-Health, a complete geolocated Maryland land parcel dataset (n=2369365) with 1317 features from calculating routing, distance, and count density of geocoded discrete SDOH resources (e.g. schools) and hazards (e.g. pollution sites), plus ancillary data (e.g. school quality). We trained a gradient boosted model (LightGBM) using all features (training sample n=1653482 parcels, testing sample n= 413371) on tract-level diabetes prevalence in CDC PLACES. We ranked top 30 features by importance. We also compared the model’s adjusted R2 and RMSE against OLS of area deprivation index (ADI) on diabetes prevalence. Results: The model had a training-sample adjusted R2 of 0.949, MAE 0.480, Spearman 0.967, and RMSE 0.667; and testing-sample adjusted R2 of 0.948, MAE 0.482, Spearman 0.967, and RMSE 0.671. Of the top 30 features by importance, 8 were related to local economic value, 8 proximity to hospitals, 6 local school performance, and 4 proximity to gun violence. The remaining 4 included proximity to food pantries, gyms, pollution sites, and FQHCs. OLS on diabetes prevalence using ADI had an adjusted R2 of 0.381 and RMSE of 2.51. Conclusion: Address-level discrete SDOH explain much of the variation in census tract diabetes prevalence in Maryland. The model’s adjusted R2 and RMSE suggested a better fit than ADI for diabetes prevalence. Our study is limited by cross-sectional analysis: future work could leverage differences in open dates for SDOH sites. Top features suggest local economic resources, geographic accessibility to healthcare, accessibility to educational quality, and exposure to violence are important potential predictors of diabetes prevalence. Disclosure S.J. Huang: None. R. Laboulaye: None. M. Bandos: None. M. Roumanos: None. O. Iyalomhe: None. A. Ratzki-Leewing: Advisory Panel; Current; Sanofi. Consultant; Current; Vertex Pharmaceuticals Incorporated, Dexcom, Inc. Research Support; Current; Sanofi. Other - Paid presentation; Ended; Abbott. Consultant; Current; Sanofi. A. Mahurkar: None. R. McCoy: None. Funding National Institutes of Health (5T32DK098107-09)
HUANG et al. (Fri,) conducted a cross-sectional in Diabetes prevalence (n=2,369,365). Address-level discrete SDOH measures vs. Area deprivation index (ADI) was evaluated on Tract-level diabetes prevalence (Adjusted R2 0.948). A machine learning model using address-level discrete social determinants of health predicted diabetes prevalence better than the area deprivation index (testing adjusted R2 0.948 vs 0.381).
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: