Abstract Traditional fixed-time traffic signal control (TSC) systems are unable to adapt to the stochastic nature of modern urban traffic, leading to excessive idling, fuel waste, and increased CO2 emissions. This paper proposes an adaptive "Self-Learning" TSC framework utilizing Deep Reinforcement Learning (DRL), specifically the Proximal Policy Optimization (PPO) algorithm. By modeling intersections as a Markov Decision Process (MDP), the agent learns optimal phase switching and duration based on real-time vehicle queue lengths and wait times. Simulations conducted in SUMO (Simulation of Urban Mobility) across a synthetic 9-intersection grid demonstrate a 33% reduction in average vehicle delay and a 21% to 27% decrease in CO2 emissions compared to conventional Webster-based controllers. This research provides a scalable architecture for smart city integration and climate-change mitigation.
Yadav et al. (Tue,) studied this question.