Accurate annual average daily traffic (AADT) estimation underpins infrastructure planning, environmental assessment, and safety analysis; however, coverage gaps persist in areas with scarce continuous counters. We propose a data-fusion framework that imputes missing AADT values by integrating crash records, clean energy infrastructure metadata, and roadway features. We evaluate a diverse set of machine learning families for tabular regression, including tree ensembles (random forest, gradient boosting), linear and regularized regression baselines, k -nearest neighbors, and shallow neural networks, and compare them against a stacked ensemble that learns to combine base-model predictions. Models are trained using an AutoML framework to standardize preprocessing, validation, and ensembling. Across traffic-only, energy-only, and hybrid feature sets, the stacked ensemble consistently achieves the lowest prediction error and remains robust across low-, medium-, and high-volume traffic regimes. On the combined feature set, the model achieves a mean absolute error (MAE) of approximately 570 vehicles/day, a root mean squared error (RMSE) of approximately 1415 vehicles/day, a mean absolute percentage error (MAPE) value of 3.73, and an R 2 value close to 0.99 under cross-validation evaluation. These results demonstrate that principled ensembling and multi-source data fusion substantially improve AADT imputation performance, particularly in data-limited settings such as uncounted or sparsely monitored roadways.
Taiwo et al. (Wed,) studied this question.