In this paper, we propose a novel, partially decentralized learning algorithm for the control of finite, multi-agent Markov Decision Process with unknown transition probabilities and reward values. One learning automaton is associated with each agent
No takes yet. Share an insight, caveat, or question.
Tilak et al. (2011) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: