This paper investigates a relatively new direction in Mul-tiagent Reinforcement Learning. Most multiagent learning techniques focus on Nash equilibria as elements of both the learning algorithm and its evaluation criteria. In contrast, we propose a multiagent learning algorithm that is optimal in the sense of finding a best-response policy, rather than in reaching an equilibrium. We present the first learning al-gorithm that is provably optimal against restricted classes of non-stationary opponents. The algorithm infers an accu-rate model of the opponent’s non-stationary strategy, and simultaneously creates a best-response policy against that strategy. Our learning algorithm works within the very gen-eral framework of n-player, general-sum stochastic games, and learns both the game structure and its associated opti-mal policy. 1.
No takes yet. Share an insight, caveat, or question.
Weinberg et al. (2004) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: