Path planning for multiple four-way shuttles in high-density warehousing is frequently hampered by efficiency-degrading conflicts, particularly head-on deadlocks. To address this challenge, this paper proposes a multi-agent reinforcement learning (MARL) framework based on Proximal Policy Optimization (PPO). The core of our approach is a novel Cooperative Avoidance Reward Mechanism (CARM), which employs a dual-component reward structure. This structure integrates a distance-guided reward to ensure efficient navigation towards targets and a cooperative avoidance reward that uses both immediate and delayed returns to incentivize implicit collaboration. This design effectively resolves conflicts and mitigates the policy instability often caused by traditional collision penalties. Experiments in a 20 × 20 grid simulation environment demonstrated that, compared to a rule-based A* and Conflict-Based Search (CBS) algorithms, the proposed method reduced the average travel distance and total time by 35.8% and 31.5%, respectively, while increasing system throughput by 49.7% and maintaining a task success rate of over 95%. Ablation studies further confirmed the critical role of CARM in achieving stable multi-agent collaboration. This work offers a scalable and efficient data-driven solution for real-time path planning in complex automated warehousing systems.
No takes yet. Share an insight, caveat, or question.
Peng et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: