In this paper we present the stability and convergence results for dynamic programming-based reinforcement learning applied to linear quadratic regulation (LQR). The specific algorithm we analyze is based on Q-learning and it is proven to converge to an optimal controller provided that the underlying system is controllable and a particular signal vector is persistently excited. This is the first convergence result for DP-based reinforcement learning algorithms for a continuous problem.
No takes yet. Share an insight, caveat, or question.
Bradtke et al. (2005) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: