PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 26, 20240 citationsOpen Access

Confident Natural Policy Gradient for Local Planning in q_-realizable Constrained MDPs

View Full Paper
TTTian TianLYLin F. YangCSCsaba Szepesvári

Key Points

Key points are not available for this paper at this time.

Abstract

The constrained Markov decision process (CMDP) framework emerges as an important reinforcement learning approach for imposing safety or other critical objectives while maximizing cumulative reward. However, the current understanding of how to learn efficiently in a CMDP environment with a potentially infinite number of states remains under investigation, particularly when function approximation is applied to the value functions. In this paper, we address the learning problem given linear function approximation with q_-realizability, where the value functions of all policies are linearly representable with a known feature map, a setting known to be more general and challenging than other linear settings. Utilizing a local-access model, we propose a novel primal-dual algorithm that, after O (poly (d) ^-3) queries, outputs with high probability a policy that strictly satisfies the constraints while nearly optimizing the value with respect to a reward function. Here, d is the feature dimension and > 0 is a given error. The algorithm relies on a carefully crafted off-policy evaluation procedure to evaluate the policy using historical data, which informs policy updates through policy gradients and conserves samples. To our knowledge, this is the first result achieving polynomial sample complexity for CMDP in the q_-realizable setting.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tian et al. (2024) studied this question.

synapsesocial.com/papers/68e634cdb6db6435875c63d5https://doi.org/10.48550/arxiv.2406.18529
Ask AI
Helpful
Bookmark
Share
View Full Paper