This paper proposes a simulation-based algorithm for optimizing the average reward in a finite-state Markov reward process that depends on a set of parameters. As a special case, the method applies to Markov decision processes where optimization takes place within a parametrized set of policies. The algorithm relies on the regenerative structure of finite-state Markov processes, involves the simulation of a single sample path, and can be implemented online. A convergence result (with probability 1) is provided.
No takes yet. Share an insight, caveat, or question.
Marbach et al. (2001) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: