Hierarchical multi-agent reinforcement learning optimizes urban traffic control, suggesting improved efficiency and sustainability outcomes.
Urbanization is intensifying congestion, emissions, and unequal mobility access in cities. This study aims to operationalize sustainability objectives—efficiency, environmental externalities, and service equity—in network-wide traffic system control. We propose SERL-H, a sustainability-aware hierarchical multi-agent reinforcement learning (MARL) controller. SERL-H separates fast intersection-level actuation from slower region-level coordination under a centralized-training decentralized-execution paradigm, and employs adaptive graph attention to capture time-varying interdependencies with bounded neighborhood communication. The learning reward explicitly balances delay/throughput, emissions/fuel, and an equity regularizer based on service dispersion across user groups. In a SUMO-based city-scale simulation with 100 signalized intersections, SERL-H reduces average delay from 45 s to 29 s and average travel time from 120 s to 88 s relative to fixed-time control, while increasing throughput and lowering total emissions (4800 kg to 3950 kg). A socio-economic assessment suggests higher annualized cost savings (e.g., $50.27 M/year to $65.91 M/year) and improved environmental quality indices. We also report, as supporting evidence, an optional sustainability-enhanced spatio-temporal graph predictor (SUT-GNN) that provides reliable short-horizon forecasts during peak-hour volatility.
No takes yet. Share an insight, caveat, or question.
Cao et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: