The policy iteration algorithm for average reward Markov decision processes with general state space | Synapse