Key points are not available for this paper at this time.
This paper presents a method called adaptive ε-greedy for better balancing between exploration and exploitation in reinforcement learning. This method is based on classic ε-greedy, which holds the value of ε statically. The solution proposed uses concepts and techniques of adaptive technology to allow controlling the value of ε during the execution. We present experiments numerically comparing the two methods for the n-armed bandit problem in a stationary environment. The adaptive ε-greedy method presents better performance as compared to the classic ε-greedy. For a nonstationary environment, we use an algorithm to detect the change point and adaptively modify the state of the agent to learn from the new rewards received. A recent related work was also evaluated in the paper.
Mignon et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: