Autonomous surface vessel (ASV) navigation research spans rule-based, optimisation-based, and learning-based paradigms, yet published results are rarely produced under identical physics, identical scenarios, or auditable scoring. This paper introduces VesselNav-Bench, an open benchmark built on a shared 3-DOF maneuvering simulator. Seven agents — a layered classical pipeline (A* + Integral Line-of-Sight guidance + predictive COLREGS avoider), Model Predictive Control (MPC), a Proximal Policy Optimisation (PPO) reinforcement learning policy trained to 10 million steps, a safety-shielded RL hybrid, an ablation baseline, and two published reactive methods (Velocity Obstacles, Artificial Potential Fields) implemented as reference baselines through the benchmark’s submission interface — are evaluated across nine seeded scenarios covering the principal COLREGS encounter types under calm and disturbed environmental conditions. Each scenario is run for 20 seeded episodes per condition (180 episodes per agent per condition; 360 across both conditions). Every decision is logged with its complete reasoning trace, enabling per-step auditability for both rule-based and learned agents. The classical pipeline and MPC top the calm leaderboard (scores 97.9 and 97.0); MPC leads under disturbances (96.3 versus 94.7 for classical), benefiting from its unified cost formulation. Safety shielding eliminates all groundings from the learned policy (11.7% → 0%) at modest efficiency cost. The RL policy improves only marginally between 4 million and 10 million training steps ( + 1.0 score points), indicating a persistent generalisation gap localised to landmass geometries underrepresented during training. Across every non-trivial agent — classical, MPC, VO, and both RL variants — COLREGS Rule 17 (stand-on) compliance is capped at 0.76, a finding with direct implications for ASV certification. Benchmark code, episode logs, and leaderboard are openly available at doi:10.5281/zenodo.20689528.
Vishnu Ravendranathan (Sat,) studied this question.