A Single Online Agent Can Efficiently Learn Mean Field Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chenyu, Chen, Xu, Di, Xuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909256325791744
author Zhang, Chenyu
Chen, Xu
Di, Xuan
author_facet Zhang, Chenyu
Chen, Xu
Di, Xuan
contents Mean field games (MFGs) are a promising framework for modeling the behavior of large-population systems. However, solving MFGs can be challenging due to the coupling of forward population evolution and backward agent dynamics. Typically, obtaining mean field Nash equilibria (MFNE) involves an iterative approach where the forward and backward processes are solved alternately, known as fixed-point iteration (FPI). This method requires fully observed population propagation and agent dynamics over the entire spatial domain, which could be impractical in some real-world scenarios. To overcome this limitation, this paper introduces a novel online single-agent model-free learning scheme, which enables a single agent to learn MFNE using online samples, without prior knowledge of the state-action space, reward function, or transition dynamics. Specifically, the agent updates its policy through the value function (Q), while simultaneously evaluating the mean field state (M), using the same batch of observations. We develop two variants of this learning scheme: off-policy and on-policy QM iteration. We prove that they efficiently approximate FPI, and a sample complexity guarantee is provided. The efficacy of our methods is confirmed by numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Single Online Agent Can Efficiently Learn Mean Field Games
Zhang, Chenyu
Chen, Xu
Di, Xuan
Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
Mean field games (MFGs) are a promising framework for modeling the behavior of large-population systems. However, solving MFGs can be challenging due to the coupling of forward population evolution and backward agent dynamics. Typically, obtaining mean field Nash equilibria (MFNE) involves an iterative approach where the forward and backward processes are solved alternately, known as fixed-point iteration (FPI). This method requires fully observed population propagation and agent dynamics over the entire spatial domain, which could be impractical in some real-world scenarios. To overcome this limitation, this paper introduces a novel online single-agent model-free learning scheme, which enables a single agent to learn MFNE using online samples, without prior knowledge of the state-action space, reward function, or transition dynamics. Specifically, the agent updates its policy through the value function (Q), while simultaneously evaluating the mean field state (M), using the same batch of observations. We develop two variants of this learning scheme: off-policy and on-policy QM iteration. We prove that they efficiently approximate FPI, and a sample complexity guarantee is provided. The efficacy of our methods is confirmed by numerical experiments.
title A Single Online Agent Can Efficiently Learn Mean Field Games
topic Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
url https://arxiv.org/abs/2405.03718