Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ocello, Antonio, Tiapkin, Daniil, Mancini, Lorenzo, Laurière, Mathieu, Moulines, Eric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037585657856
author Ocello, Antonio
Tiapkin, Daniil
Mancini, Lorenzo
Laurière, Mathieu
Moulines, Eric
author_facet Ocello, Antonio
Tiapkin, Daniil
Mancini, Lorenzo
Laurière, Mathieu
Moulines, Eric
contents We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite state-action spaces. Building on the well-established performance of TRPO in the reinforcement learning (RL) setting, we extend its methodology to the MFG framework, leveraging its stability and robustness in policy optimization. Under standard assumptions in the MFG literature, we provide a rigorous analysis of MF-TRPO, establishing theoretical guarantees on its convergence. Our results cover both the exact formulation of the algorithm and its sample-based counterpart, where we derive high-probability guarantees and finite sample complexity. This work advances MFG optimization by bridging RL techniques with mean-field decision-making, offering a theoretically grounded approach to solving complex multi-agent problems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
Ocello, Antonio
Tiapkin, Daniil
Mancini, Lorenzo
Laurière, Mathieu
Moulines, Eric
Machine Learning
Statistics Theory
We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite state-action spaces. Building on the well-established performance of TRPO in the reinforcement learning (RL) setting, we extend its methodology to the MFG framework, leveraging its stability and robustness in policy optimization. Under standard assumptions in the MFG literature, we provide a rigorous analysis of MF-TRPO, establishing theoretical guarantees on its convergence. Our results cover both the exact formulation of the algorithm and its sample-based counterpart, where we derive high-probability guarantees and finite sample complexity. This work advances MFG optimization by bridging RL techniques with mean-field decision-making, offering a theoretically grounded approach to solving complex multi-agent problems.
title Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2505.22781