We propose a new constrained Markov decision process framework with risk-type constraints. The risk metric we use is Conditional Value-at-Risk (CVaR), which is gaining popularity in finance. It is a conditional expectation but the conditioning is defined in terms of the level of the tail probability. We propose an iterative offline algorithm to find the risk-contrained optimal control policy. A two time-scale stochastic approximation-inspired `learning' variant is also sketched, and its convergence proved to the optimal risk-constrained policy.
No takes yet. Share an insight, caveat, or question.
Borkar et al. (2014) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: