PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 8, 20203 citationsOpen Access

Hallucinating Value: A Pitfall of Dyna-style Planning with Imperfect Environment Models

TJTaher JafferjeeEIEhsan ImaniETErin J. Talvitie

Key Points

Key points are not available for this paper at this time.

Abstract

Dyna-style reinforcement learning (RL) agents improve sample efficiency over-free RL agents by updating the value function with simulated experience by an environment model. However, it is often difficult to learn models of environment dynamics, and even small errors may result in of Dyna agents. In this paper, we investigate one type of model error: states. These are states generated by the model, but that are not states of the environment. We present the Hallucinated Value Hypothesis (HVH): updating values of real states towards values of hallucinated states in misleading state-action values which adversely affect the control. We discuss and evaluate four Dyna variants; three which update real toward simulated -- and therefore potentially hallucinated -- states and which does not. The experimental results provide evidence for the HVH thus a fruitful direction toward developing Dyna algorithms robust to error.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jafferjee et al. (2020) studied this question.

synapsesocial.com/papers/6a170e68f96f07bf256b8b4ahttps://doi.org/10.48550/arxiv.2006.04363
Ask AI
Helpful
Bookmark
Share
View Full Paper