Key points are not available for this paper at this time.
Dynamic routing, modulation, and spectrum assignment (RMSA) in elastic optical networks (EONs) requires joint optimization considering complex physical layer impairments. While deep reinforcement learning (DRL) has shown promise for RMSA, existing methods face two fundamental limitations: (i) rigid distance-adaptive modulation rules that underutilize spectrum resources and (ii) value estimation bias in continuing tasks that prevents convergence to optimal policies. This paper proposes a physical layer-aware DRL framework that addresses both limitations. First, we incorporate reward centering to eliminate value estimation bias in continuing tasks, enabling the agent to distinguish fine-grained policy differences. Second, the framework enables autonomous joint optimization of routing and modulation selection, removing reliance on distance-based rules. Simulations on NSFNET and COST239 demonstrate two key results: (i) reward centering reduces service blocking probability by 16% compared to standard DRL under identical constraints, and (ii) autonomous modulation selection reduces blocking by up to 77% in high-load regimes where distance-adaptive methods saturate at approximately 16%. Physical layer analysis reveals that performance gains are achieved by operating closer to transmission limits, with the average GSNR margin reduced from 7.1 to 2.7 dB.
Wang et al. (Mon,) studied this question.