July 4, 2024Open Access

A Two-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games

Key Points

Key points are not available for this paper at this time.

Abstract

An interesting iterative procedure is proposed to solve a two-player zero-sum Markov games. First this problem is expressed as a min-max Markov game. Next, a two-step Q-learning algorithm for solving Markov decision problem (MDP) is suitably modified to solve this Markov game. Under a suitable assumption, the boundedness of the proposed iterates is obtained theoretically. Using results from stochastic approximation, the almost sure convergence of the proposed two-step minimax Q-learning is obtained theoretically. More specifically, the proposed algorithm converges to the game theoretic optimal value with probability one, when the model information is not known. Numerical simulation authenticate that the proposed algorithm is effective and easy to implement.

Read Full Paperexternally

Bookmark

View Full Paper

Cite This Study

Shreyas et al. (Thu,) studied this question.

synapsesocial.com/papers/68e616ccb6db6435875a9aa3 https://doi.org/https://doi.org/10.48550/arxiv.2407.04240

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

Bookmark

View Full Paper