We consider a discrete time Markov decision process, where the objectives are linear combinations of standard discounted rewards, each with a different discount factor. We describe several applications that motivate the recent interest in these criteria. For the special case where a standard discounted cost is to be minimized, subject to a constraint on another standard discounted cost but with a different discount factor, we provide an implementable algorithm for computing an optimal policy.
No takes yet. Share an insight, caveat, or question.
Feinberg et al. (1999) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: