This paper considers a distributed reinforcement learning problem in which a network of multiple agents aim to cooperatively maximize the globally averaged return through communication with only local neighbors. A randomized communication-efficient multi-agent actor-critic algorithm is proposed for possibly unidirectional communication relationships depicted by a directed graph. It is shown that the algorithm can solve the problem for strongly connected graphs by allowing each agent to transmit only two scalar-valued variables at one time.
No takes yet. Share an insight, caveat, or question.
Lin et al. (2019) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: