PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 20240 citationsOpen Access

Finite Time Analysis of Temporal Difference Learning for Mean-Variance in a Discounted MDP

View Full Paper
TSTejaram SangadiLPL. A. PrashanthIndian Institute of Technology MadrasKJKrishna JagannathanIndian Institute of Technology Madras

Key Points

Key points are not available for this paper at this time.

Abstract

Motivated by risk-sensitive reinforcement learning scenarios, we consider the problem of policy evaluation for variance in a discounted reward Markov decision process (MDP). For this problem, a temporal difference (TD) type learning algorithm with linear function approximation (LFA) exists in the literature, though only asymptotic guarantees are available for this algorithm. We derive finite sample bounds that hold (i) in the mean-squared sense; and (ii) with high probability, when tail iterate averaging is employed with/without regularization. Our bounds exhibit exponential decay for the initial error, while the overall bound is O (1/t), where t is the number of update iterations of the TD algorithm. Further, the bound for the regularized TD variant is for a universal step size. Our bounds open avenues for analysis of actor-critic algorithms for mean-variance optimization in a discounted MDP.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sangadi et al. (2024) studied this question.

synapsesocial.com/papers/68e651c6b6db6435875e24b6https://doi.org/10.48550/arxiv.2406.07892
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1High-Probability Sample Complexities for Policy Evaluation With Linear Function Approximation2024
  2. 2An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks2024
  3. 3Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features2025
  4. 4The surprising efficiency of temporal difference learning for rare event prediction2024
  5. 5A Simple Finite-Time Analysis of TD Learning with Linear Function Approximation2024