PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 1, 201923 citations

Approximate information state for partially observed systems

View Full Paper
JSJayakumar SubramanianAMAditya Mahajan

Key Points

Key points are not available for this paper at this time.

Abstract

The standard approach for modeling partially observed systems is to model them as partially observable Markov decision processes (POMDPs) and obtain a dynamic program in terms of a belief state. The belief state formulation works well for planning but is not ideal for online reinforcement learning because the belief state depends on the model and, as such, is not observable when the model is unknown.In this paper, we present an alternative notion of an information state for obtaining a dynamic program in partially observed models. In particular, an information state is a sufficient statistic for the current reward which evolves in a controlled Markov manner. We show that such an information state leads to a dynamic programming decomposition. Then we present a notion of an approximate information state and present an approximate dynamic program based on the approximate information state. Approximate information state is defined in terms of properties that can be estimated using sampled trajectories. Therefore, they provide a constructive method for reinforcement learning in partially observed systems. We present one such construction and show that it performs better than the state of the art for three benchmark models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Subramanian et al. (2019) studied this question.

synapsesocial.com/papers/6a0eeb6e9df4132b62f9ceeahttps://doi.org/10.1109/cdc40024.2019.9029898
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Metrics for Finite Markov Decision Processes2012 · 54 citations
  2. 2Simple statistical gradient-following algorithms for connectionist reinforcement learning1992 · 7,522 citations
  3. 3Optimal Transport: Old and New2013 · 3,935 citations
  4. 4Discrete-Time Markov Control Processes: Basic Optimality Criteria.1996 · 193 citations