Simulation of single-agent RL improves bus headway and occupancy performance, suggesting a viable alternative to MARL.
Bus bunching remains a critical challenge in urban transit systems, primarily driven by the stochastic nature of traffic conditions and passenger demand. Recently a popular method to address this issue is multi-agent reinforcement learning (MARL) applied in an idealized loop-line environment. However, such method generally suffers from high computational costs and sample inefficiency. Moreover, they often fail to capture the dynamics of realistic bus systems, which are typically governed by heterogeneous trip line and variable fleet sizes. In this study, we propose a robust single-agent reinforcement learning (RL) framework for bus holding control in a bidirectional timetabled bus line, explicitly designed to circumvent the data imbalance and convergence issues associated with MARL. Our key contribution focuses on transforming the inherently multi-agent problem into a single-agent formulation by explicitly encoding categorical identifiers—such as vehicle, station, and trip IDs alongside traditional numerical features (e.g., headway, occupancy and segment velocity) together as the state representation. This feature space augmentation enables the single-agent to operate effectively in a higher-dimensional space, analogous to projecting linearly inseparable inputs into a higher-dimensional space to achieve separability. Additionally, we introduce a structured ”ridge-shaped” reward function that incentivizes the alignment with both uniform headways and scheduled departure intervals. Compared to the other three benchmark methods, our proposed RL algorithm achieves significantly more stable and higher rewards (-430k comparing with -530k) under stochastic passenger demands and inter-station travel time. These experimental results suggest that the proposed single-agent RL approach, informed by categorical identifiers in the state representation and realistic schedule-aware design in the ridge-shaped reward function, can effectively mitigate bus bunching in non-loop settings. This paradigm offers a robust and scalable alternative to those conventional MARL-based control frameworks, particularly in environments where agent-specific experience distributions are inherently imbalanced.
No takes yet. Share an insight, caveat, or question.
Yifan Zhang (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: