Autonomous navigation policies trained via reinforcement learning degrade silently under distribution shift, creating safety hazards with no warning signal. We propose Uncertainty-Gated Selective Deployment (UGSD), a principled framework that combines runtime epistemic uncertainty estimation with an adaptive routing policy to decide—on a per-episode basis—whether to navigate autonomously or request human intervention. UGSD introduces three technical contributions: (1) a distribution shift decomposition protocol that factorizes OOD conditions into sensor-degradation and structural-novelty components, enabling targeted uncertainty analysis; (2) an adaptive threshold routing algorithm that optimizes the autonomy–safety trade-off by selecting deployment thresholds from calibration data; and (3) a layer-wise uncertainty propagation analysis explaining why MC-Dropout (T=20) achieves AUROC 0.967 for failure prediction—outperforming a 5× more expensive deep ensemble (AUROC 0.721)—on compound distribution shift. We evaluate across a controlled 4-environment shift spectrum in high-fidelity simulation with a ROS 2/Gazebo TurtleBot3 platform. UGSD raises effective success rate from 37.5% to 98.0% on the autonomously-handled subset at 75% human routing burden. Post-hoc temperature scaling reduces ECE from 0.47 to 0.08 on layout-shifted environments. A decomposition analysis reveals that structural novelty produces 2.7–3.2× more epistemic uncertainty than sensor degradation alone, providing actionable guidance for deployment risk assessment. The system is implemented as a ROS 2 Humble package validated in Gazebo simulation with a TurtleBot3 platform.
Newton Adhikari (Fri,) studied this question.