We show how an ensemble of Q^*-functions can be leveraged for more effective exploration in deep reinforcement learning. We build on well established algorithms from the bandit setting, and adapt them to the Q-learning setting. We propose an exploration strategy based on upper-confidence bounds (UCB). Our experiments show significant gains on the Atari benchmark.
No takes yet. Share an insight, caveat, or question.
Chen et al. (2017) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: