Urban road networks lose a large amount of time and fuel to congestion at signalised intersections, and conventional fixed-time controllers cannot react to the changing and uneven traffic demand seen in practice. Reinforcement learning (RL), and in particular deep reinforcement learning (DRL), has been proposed as a way to learn signal-timing policies directly from interaction with traffic. This paper proposes a simulation-based framework for adaptive traffic signal control at a single four-way intersection, built on the SUMO microscopic traffic simulator. Two DRL agents, a Deep Q-Network (DQN) and Proximal Policy Optimization (PPO), are formulated as Markov decision processes with a queue- and waiting-time-based state, a phase-selection action space, and a reward based on the change in cumulative vehicle waiting time. The framework defines a comparative evaluation protocol in which both agents are measured against fixed-time, vehicle-actuated and max-pressure controllers across five demand scenarios, using average delay, queue length, stops, and throughput over multiple random seeds. This is a proposal paper: the framework is specified in full, but no experiments have been run, so no results are reported. The paper also states the expected outcomes, limitations, and directions for implementation.
No takes yet. Share an insight, caveat, or question.
Swarnal Deshmukh (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: