PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 20240 citationsOpen Access

Regularized Q-learning through Robust Averaging

View Full Paper
PSPeter Schmitt-FörsterTSTobias Sutter

Key Points

Key points are not available for this paper at this time.

Abstract

We propose a new Q-learning variant, called 2RA Q-learning, that addresses some weaknesses of existing Q-learning methods in a principled manner. One such weakness is an underlying estimation bias which cannot be controlled and often results in poor performance. We propose a distributionally robust estimator for the maximum expected value term, which allows us to precisely control the level of estimation bias introduced. The distributionally robust estimator admits a closed-form solution such that the proposed algorithm has a computational cost per iteration comparable to Watkins' Q-learning. For the tabular case, we show that 2RA Q-learning converges to the optimal policy and analyze its asymptotic mean-squared error. Lastly, we conduct numerical experiments for various settings, which corroborate our theoretical findings and indicate that 2RA Q-learning often performs better than existing methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schmitt-Förster et al. (2024) studied this question.

synapsesocial.com/papers/68e6bbccb6db64358763c50fhttps://doi.org/10.48550/arxiv.2405.02201
Ask AI
Helpful
Bookmark
Share
View Full Paper