Comparative modeling study demonstrates causal feature selection improves epidemic forecasting accuracy and interpretability, indicating strong utility for guiding public health interventions.
Forecasting requires identifying the sources of information that are most relevant for quantifying future projections. This is important because including unnecessary or misleading information can reduce forecasting performance and make the results more difficult to interpret. This paper examines the role of causal feature selection in time series forecasting and argues that it should be incorporated as a component of machine learning forecasting frameworks. Several widely used feature selection methods are compared, including Principal Component Analysis, SHapley Additive exPlanations, and subsets of lagged time series data, in the context of modeling the spread of COVID-19 in Ontario, Canada using epidemiological and mobility data. Forecasting performance is evaluated using R-squared, Mean Absolute Error, Root Mean Squared Error, Mean Absolute Percentage Error, and forecast bias. The results indicate that machine learning models trained on features selected with the estimated direct causal relations consistently outperform models trained on association-based selection. Beyond forecasting performance gains, causal feature selection provides explainable models by identifying actionable drivers that can be targeted to influence disease dynamics – in contrast to statistical approaches that capture associational patterns – thereby enhancing the suitability of such models for decision making and other optimization contexts where interventions impact future outcomes.
No takes yet. Share an insight, caveat, or question.
Mossop et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: