PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 20240 citationsOpen Access

Orthogonalized Estimation of Difference of Q-functions

View Full Paper
AZAngela Zhou

Key Points

Key points are not available for this paper at this time.

Abstract

Offline reinforcement learning is important in many settings with available observational data but the inability to deploy new policies online due to safety, cost, and other concerns. Many recent advances in causal inference and machine learning target estimation of causal contrast functions such as CATE, which is sufficient for optimizing decisions and can adapt to potentially smoother structure. We develop a dynamic generalization of the R-learner (Nie and Wager 2021, Lewis and Syrgkanis 2021) for estimating and optimizing the difference of Q^-functions, Q^ (s, 1) -Q^ (s, 0) (which can be used to optimize multiple-valued actions). We leverage orthogonal estimation to improve convergence rates in the presence of slower nuisance estimation rates and prove consistency of policy optimization under a margin condition. The method can leverage black-box nuisance estimators of the Q-function and behavior policy to target estimation of a more structured Q-function contrast.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Angela Zhou (2024) studied this question.

synapsesocial.com/papers/68e651cbb6db6435875e2784https://doi.org/10.48550/arxiv.2406.08697
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Conservative Q-Learning for Offline Reinforcement Learning2020 · 538 citations
  2. 2Isolated Q-learning for Offline Reinforcement Learning2024
  3. 3A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning2024 · 17 citations
  4. 4Strategically Conservative Q-Learning2024
  5. 5Equivariant Offline Reinforcement Learning2024