On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Labbi, Safwan, Mangold, Paul, Tiapkin, Daniil, Moulines, Eric
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911559030145024
author Labbi, Safwan
Mangold, Paul
Tiapkin, Daniil
Moulines, Eric
author_facet Labbi, Safwan
Mangold, Paul
Tiapkin, Daniil
Moulines, Eric
contents We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient with explicit constants, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the Łojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23459
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
Labbi, Safwan
Mangold, Paul
Tiapkin, Daniil
Moulines, Eric
Machine Learning
We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient with explicit constants, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the Łojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies.
title On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
topic Machine Learning
url https://arxiv.org/abs/2505.23459