Learning Mean-Field Games through Mean-Field Actor-Critic Flow

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Mo, Zhou, Haosheng, Hu, Ruimeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914110990450688
author Zhou, Mo
Zhou, Haosheng
Hu, Ruimeng
author_facet Zhou, Mo
Zhou, Haosheng
Hu, Ruimeng
contents We propose the Mean-Field Actor-Critic (MFAC) flow, a continuous-time learning dynamics for solving mean-field games (MFGs), combining techniques from reinforcement learning and optimal transport. The MFAC framework jointly evolves the control (actor), value function (critic), and distribution components through coupled gradient-based updates governed by partial differential equations (PDEs). A central innovation is the Optimal Transport Geodesic Picard (OTGP) flow, which drives the distribution toward equilibrium along Wasserstein-2 geodesics. We conduct a rigorous convergence analysis using Lyapunov functionals and establish global exponential convergence of the MFAC flow under a suitable timescale. Our results highlight the algorithmic interplay among actor, critic, and distribution components. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the MFAC framework in computing MFG equilibria.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Mean-Field Games through Mean-Field Actor-Critic Flow
Zhou, Mo
Zhou, Haosheng
Hu, Ruimeng
Optimization and Control
Machine Learning
35Q89, 49N80
We propose the Mean-Field Actor-Critic (MFAC) flow, a continuous-time learning dynamics for solving mean-field games (MFGs), combining techniques from reinforcement learning and optimal transport. The MFAC framework jointly evolves the control (actor), value function (critic), and distribution components through coupled gradient-based updates governed by partial differential equations (PDEs). A central innovation is the Optimal Transport Geodesic Picard (OTGP) flow, which drives the distribution toward equilibrium along Wasserstein-2 geodesics. We conduct a rigorous convergence analysis using Lyapunov functionals and establish global exponential convergence of the MFAC flow under a suitable timescale. Our results highlight the algorithmic interplay among actor, critic, and distribution components. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the MFAC framework in computing MFG equilibria.
title Learning Mean-Field Games through Mean-Field Actor-Critic Flow
topic Optimization and Control
Machine Learning
35Q89, 49N80
url https://arxiv.org/abs/2510.12180