PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 20240 citationsOpen Access

An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks

View Full Paper
ZKZhifa KeZWZaiwen WenJZJunyu Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

Temporal difference (TD) learning algorithms with neural network function parameterization have well-established empirical success in many practical large-scale reinforcement learning tasks. However, theoretical understanding of these algorithms remains challenging due to the nonlinearity of the action-value approximation. In this paper, we develop an improved non-asymptotic analysis of the neural TD method with a general L-layer neural network. New proof techniques are developed and an improved new O (^-1) sample complexity is derived. To our best knowledge, this is the first finite-time analysis of neural TD that achieves an O (^-1) complexity under the Markovian sampling, as opposed to the best known O (^-2) complexity in the existing literature.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ke et al. (2024) studied this question.

synapsesocial.com/papers/68e6b4c2b6db6435876358e2https://doi.org/10.48550/arxiv.2405.04017
Ask AI
Helpful
Bookmark
Share
View Full Paper