This study introduces a novel dam operation framework using Deep Reinforcement Learning (DRL) that operates without any inflow forecasting. Focusing on Japan's Hiyoshi Dam, a Deep Q-Network (DQN) model was trained on historical flood events using only observed data, such as inflow, outflow, reservoir storage, and local meteorological conditions, as inputs. The agent was evaluated across both trained and unseen flood scenarios to assess generalization and robustness. Two distinct reward functions were designed to explore their impact on policy behavior: one encouraging proactive discharge actions and the other favoring smoother operations. The DQN models achieved up to 99% of the cumulative reward compared to dynamic programming benchmarks, even in untrained flood conditions. Sensitivity analyses demonstrated the model’s adaptability to different flood onset times, initial storage levels, and input perturbations, highlighting both strengths and limitations in operational realism. These findings underscore the feasibility of DRL for real-time dam operation without relying on uncertain forecasts and offer new insights into reward design, policy generalization, and hydrologic robustness under climate-induced extremes.
Kim et al. (Thu,) studied this question.