On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911559030145024 |
|---|---|
| author | Labbi, Safwan Mangold, Paul Tiapkin, Daniil Moulines, Eric |
| author_facet | Labbi, Safwan Mangold, Paul Tiapkin, Daniil Moulines, Eric |
| contents | We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient with explicit constants, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the Łojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_23459 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments Labbi, Safwan Mangold, Paul Tiapkin, Daniil Moulines, Eric Machine Learning We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient with explicit constants, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the Łojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies. |
| title | On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.23459 |