Key points are not available for this paper at this time.
Air pollution is a major factor influencing hospital admissions worldwide, highlighting the need for robust predictive tools to support healthcare planning and public health measures. Machine learning (ML) has been widely employed to simulate the intricate relationships between pollution and health outcomes. This paper examines publications indexed in the Scopus database, from 2010 to 2024 focusing on using ML techniques to forecast outcomes related to air pollution and hospital admissions. A bibliometric study of the 89 identified papers was also conducted to determine dominant research themes, commonly employed methodologies, and the geographical distribution of publications. The results indicate that research activity increased notably after 2020, with the United States of America, China, and Brazil contributing the highest number of publications. Moreover, the findings indicate that approximately 83% of the reviewed research applied predictive models appropriately, suggesting that ML techniques can effectively forecast healthcare outcomes. Random Forest was the most frequently used method (33 studies), followed by Neural Networks (18 studies). Extreme Gradient Boosting (XGBoost) algorithm, although less frequent, showed the highest reported accuracy, with values ranging from 87% to 95%. The most studied pollutants were particulate matter (PM2.5), nitrogen dioxide (NO2), and coarse particulate matter (PM10). Demographic and meteorological data were the most frequently used complementary (71% and 65%, respectively), followed by temporal (46%) and socioeconomic factors (20%). The combination of several variable categories not only enhanced understanding of how environmental exposure affects health outcomes but also improved the accuracy and reliability of the reviewed ML models.
Aval et al. (Tue,) studied this question.