Offline Two-Player Zero-Sum Markov Games with KL Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Claire, Zhang, Yuheng, Liu, Xinyu, Xie, Zixuan, Liu, Shuze Daniel, Jiang, Nan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913122392997888
author Chen, Claire
Zhang, Yuheng
Liu, Xinyu
Xie, Zixuan
Liu, Shuze Daniel
Jiang, Nan
author_facet Chen, Claire
Zhang, Yuheng
Liu, Xinyu
Xie, Zixuan
Liu, Shuze Daniel
Jiang, Nan
contents We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to stabilize learning and guarantee convergence. We first introduce Regularized Offline Sequential Equilibrium (ROSE), a theoretical framework that achieves a fast $\widetilde{\mathcal{O}}(1/n)$ convergence rate under \textit{unilateral concentrability}, improving over the standard $\widetilde{\mathcal{O}}(1/\sqrt{n})$ rates in unregularized settings. We then propose Sequential Offline Self-play Mirror Descent (SOS-MD), a practical model-free algorithm based on least-squares value estimation and iterative self-play updates. We prove that the last iterate of SOS-MD attains the same $\widetilde{\mathcal{O}}(1/n)$ statistical rate up to a vanishing optimization error of order $\widetilde{\mathcal{O}}(1/\sqrt{T})$ in the number of self-play iterations $T$.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13025
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Offline Two-Player Zero-Sum Markov Games with KL Regularization
Chen, Claire
Zhang, Yuheng
Liu, Xinyu
Xie, Zixuan
Liu, Shuze Daniel
Jiang, Nan
Machine Learning
Computer Science and Game Theory
We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to stabilize learning and guarantee convergence. We first introduce Regularized Offline Sequential Equilibrium (ROSE), a theoretical framework that achieves a fast $\widetilde{\mathcal{O}}(1/n)$ convergence rate under \textit{unilateral concentrability}, improving over the standard $\widetilde{\mathcal{O}}(1/\sqrt{n})$ rates in unregularized settings. We then propose Sequential Offline Self-play Mirror Descent (SOS-MD), a practical model-free algorithm based on least-squares value estimation and iterative self-play updates. We prove that the last iterate of SOS-MD attains the same $\widetilde{\mathcal{O}}(1/n)$ statistical rate up to a vanishing optimization error of order $\widetilde{\mathcal{O}}(1/\sqrt{T})$ in the number of self-play iterations $T$.
title Offline Two-Player Zero-Sum Markov Games with KL Regularization
topic Machine Learning
Computer Science and Game Theory
url https://arxiv.org/abs/2605.13025