PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026Operations Research2 citations

Beyond Discounted Returns: Robust Markov Decision Processes with Average and Blackwell Optimality

View Full Paper
MPMarek PetrikMPMarek PetrikNVNicolas Vieille

Key Points

  • This research aims to investigate robust Markov decision processes (RMDPs) for average and Blackwell optimality criteria beyond discounted returns.
  • Analyzed average optimal policies in sa-rectangular and s-rectangular RMDPs.
  • Examined the existence and nature of stationary policies.
  • Studied Blackwell optimality and provided conditions for existence.
  • Developed algorithms to compute optimal average returns.
  • Explored connections between RMDPs and stochastic games.
  • Stationary and deterministic average optimal policies exist for sa-rectangular RMDPs.
  • Average optimal policies may not exist for s-rectangular RMDPs.
  • Approximately Blackwell optimal policies always exist for sa-rectangular RMDPs.
  • Provided a sufficient condition for the existence of Blackwell optimal policies.
  • Emphasized the advantages of distance-based sa-rectangular models over s-rectangular models.

Abstract

Novel Insights on Robust Markov decision Processes with Average Reward and Blackwell Optimality Criteria Robust Markov decision processes (RMDPs) have been studied extensively when the objective is the discounted return, but little is known for average optimality and Blackwell optimality. We show that average optimal policies can be chosen stationary and deterministic for sa-rectangular RMDPs, but perhaps surprisingly, we show that for s-rectangular RMDPs average optimal policies may not exist, and if they do exist, they may not be stationary. We also study Blackwell optimality for sa-rectangular RMDPs, where we show that approximately Blackwell optimal policies always exist, although exact Blackwell optimal policies may not exist. We provide a general sufficient condition for their existence. We then discuss the connection between average and Blackwell optimality, and we describe several algorithms to compute the optimal average return. Interestingly, our approach leverages the connections between RMDPs and stochastic games. Overall, our paper emphasizes the superior practical properties of distance-based sa-rectangular models over s-rectangular models for average and Blackwell optimality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Petrik et al. (2026) studied this question.

synapsesocial.com/papers/69aa70e7531e4c4a9ff5b116https://doi.org/10.1287/opre.2023.0694
Ask AI
Helpful
Bookmark
Share
View Full Paper