The paper deals with continuous time Markov decision processes on a fairly general state space. The rewards are continuously discounted at rate \α > 0. A set of conditions is shown to be necessary and sufficient for a policy to be optimal. For the special case of time independent reward function and under the assumption that the action space is finite a policy improvement algorithm is proposed and its convergence to an optimal policy is proved.
No takes yet. Share an insight, caveat, or question.
Bharat T. Doshi (1976) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: