No-regret learning in harmonic games: Extrapolation in the face of conflicting interests

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Legacci, Davide, Mertikopoulos, Panayotis, Papadimitriou, Christos H., Piliouras, Georgios, Pradelski, Bary S. R.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913629442408448
author Legacci, Davide
Mertikopoulos, Panayotis
Papadimitriou, Christos H.
Piliouras, Georgios
Pradelski, Bary S. R.
author_facet Legacci, Davide
Mertikopoulos, Panayotis
Papadimitriou, Christos H.
Piliouras, Georgios
Pradelski, Bary S. R.
contents The long-run behavior of multi-agent learning - and, in particular, no-regret learning - is relatively well-understood in potential games, where players have aligned interests. By contrast, in harmonic games - the strategic counterpart of potential games, where players have conflicting interests - very little is known outside the narrow subclass of 2-player zero-sum games with a fully-mixed equilibrium. Our paper seeks to partially fill this gap by focusing on the full class of (generalized) harmonic games and examining the convergence properties of follow-the-regularized-leader (FTRL), the most widely studied class of no-regret learning schemes. As a first result, we show that the continuous-time dynamics of FTRL are Poincaré recurrent, that is, they return arbitrarily close to their starting point infinitely often, and hence fail to converge. In discrete time, the standard, "vanilla" implementation of FTRL may lead to even worse outcomes, eventually trapping the players in a perpetual cycle of best-responses. However, if FTRL is augmented with a suitable extrapolation step - which includes as special cases the optimistic and mirror-prox variants of FTRL - we show that learning converges to a Nash equilibrium from any initial condition, and all players are guaranteed at most O(1) regret. These results provide an in-depth understanding of no-regret learning in harmonic games, nesting prior work on 2-player zero-sum games, and showing at a high level that harmonic games are the canonical complement of potential games, not only from a strategic, but also from a dynamic viewpoint.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle No-regret learning in harmonic games: Extrapolation in the face of conflicting interests
Legacci, Davide
Mertikopoulos, Panayotis
Papadimitriou, Christos H.
Piliouras, Georgios
Pradelski, Bary S. R.
Computer Science and Game Theory
Machine Learning
Multiagent Systems
Optimization and Control
Primary 91A10, 91A26, secondary 68Q32, 68T02
The long-run behavior of multi-agent learning - and, in particular, no-regret learning - is relatively well-understood in potential games, where players have aligned interests. By contrast, in harmonic games - the strategic counterpart of potential games, where players have conflicting interests - very little is known outside the narrow subclass of 2-player zero-sum games with a fully-mixed equilibrium. Our paper seeks to partially fill this gap by focusing on the full class of (generalized) harmonic games and examining the convergence properties of follow-the-regularized-leader (FTRL), the most widely studied class of no-regret learning schemes. As a first result, we show that the continuous-time dynamics of FTRL are Poincaré recurrent, that is, they return arbitrarily close to their starting point infinitely often, and hence fail to converge. In discrete time, the standard, "vanilla" implementation of FTRL may lead to even worse outcomes, eventually trapping the players in a perpetual cycle of best-responses. However, if FTRL is augmented with a suitable extrapolation step - which includes as special cases the optimistic and mirror-prox variants of FTRL - we show that learning converges to a Nash equilibrium from any initial condition, and all players are guaranteed at most O(1) regret. These results provide an in-depth understanding of no-regret learning in harmonic games, nesting prior work on 2-player zero-sum games, and showing at a high level that harmonic games are the canonical complement of potential games, not only from a strategic, but also from a dynamic viewpoint.
title No-regret learning in harmonic games: Extrapolation in the face of conflicting interests
topic Computer Science and Game Theory
Machine Learning
Multiagent Systems
Optimization and Control
Primary 91A10, 91A26, secondary 68Q32, 68T02
url https://arxiv.org/abs/2412.20203