Simulation shows improved convergence rates in discrete-time linear-quadratic optimal control, suggesting better performance than policy iteration.
The paper demonstrates via simulation that the well-known value iteration algorithm of reinforcement learning for the discrete-time linear-quadratic optimal control problem converges very slowly—at most linearly. Despite its slow convergence, the value iteration algorithm still converges even when the initial feedback gain is several orders of magnitude away from the optimal one, and even when the initial feedback gain is not stabilizing, as demonstrated by an example. It is known that the convergence rate of the corresponding policy iteration algorithm is quadratic, assuming the initial feedback gain is stabilizing. We show that the convergence speed of the value iteration algorithm can also be made quadratic by applying ideas from the doubling algorithm used for solving the algebraic Riccati equation. We precisely state a condition required for convergence of the value iteration algorithm, which turns out to be milder than the corresponding condition for the policy iteration algorithm. In addition, we show that the newly proposed value iteration algorithm requires less computational effort than the policy iteration algorithm. With these improvements and observations, we revitalize the value iteration algorithm and demonstrate its superiority over the policy iteration algorithm.
No takes yet. Share an insight, caveat, or question.
Xu et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: