On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Labbi, Safwan, Mangold, Paul, Tiapkin, Daniil, Moulines, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
by: Labbi, Safwan, et al.
Published: (2024)
by: Labbi, Safwan, et al.
Published: (2024)
Refined Analysis of Entropy-Regularized Actor-Critic
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
by: Mangold, Paul, et al.
Published: (2024)
by: Mangold, Paul, et al.
Published: (2024)
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
by: Ocello, Antonio, et al.
Published: (2025)
by: Ocello, Antonio, et al.
Published: (2025)
Joint Channel Selection using FedDRL in V2X
by: Mancini, Lorenzo, et al.
Published: (2024)
by: Mancini, Lorenzo, et al.
Published: (2024)
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
by: Samsonov, Sergey, et al.
Published: (2023)
by: Samsonov, Sergey, et al.
Published: (2023)
Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Refined Analysis of Federated Averaging and Federated Richardson-Romberg
by: Mangold, Paul, et al.
Published: (2024)
by: Mangold, Paul, et al.
Published: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Model-free Posterior Sampling via Learning Rate Randomization
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
by: Levin, Ilya, et al.
Published: (2026)
by: Levin, Ilya, et al.
Published: (2026)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
by: Tiapkin, Daniil, et al.
Published: (2024)
by: Tiapkin, Daniil, et al.
Published: (2024)
Revisiting Non-Acyclic GFlowNets in Discrete Environments
by: Morozov, Nikita, et al.
Published: (2025)
by: Morozov, Nikita, et al.
Published: (2025)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
by: Gritsaev, Timofei, et al.
Published: (2024)
by: Gritsaev, Timofei, et al.
Published: (2024)
Proximal Point Nash Learning from Human Feedback
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Demonstration-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
by: Klein, Sara, et al.
Published: (2023)
by: Klein, Sara, et al.
Published: (2023)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
On the Convergence Rates of Federated Q-Learning across Heterogeneous Environments
by: Wang, Leo Muxing, et al.
Published: (2024)
by: Wang, Leo Muxing, et al.
Published: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Incentivized Learning in Principal-Agent Bandit Games
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Learning Shortest Paths with Generative Flow Networks
by: Morozov, Nikita, et al.
Published: (2026)
by: Morozov, Nikita, et al.
Published: (2026)
Adaptive Set-Mass Calibration with Conformal Prediction
by: Kazantsev, Daniil, et al.
Published: (2025)
by: Kazantsev, Daniil, et al.
Published: (2025)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
by: Goel, Gautam, et al.
Published: (2026)
by: Goel, Gautam, et al.
Published: (2026)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Accelerated Policy Gradient: On the Convergence Rates of the Nesterov Momentum for Reinforcement Learning
by: Chen, Yen-Ju, et al.
Published: (2023)
by: Chen, Yen-Ju, et al.
Published: (2023)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Client Selection for Federated Policy Optimization with Environment Heterogeneity
by: Xie, Zhijie, et al.
Published: (2023)
by: Xie, Zhijie, et al.
Published: (2023)
Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch
by: Zhang, Shangtong, et al.
Published: (2021)
by: Zhang, Shangtong, et al.
Published: (2021)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
by: Zhu, Zhaoyu, et al.
Published: (2026)
by: Zhu, Zhaoyu, et al.
Published: (2026)
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
Efficient Conformal Prediction under Data Heterogeneity
by: Plassier, Vincent, et al.
Published: (2023)
by: Plassier, Vincent, et al.
Published: (2023)
Similar Items
-
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
by: Labbi, Safwan, et al.
Published: (2026) -
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
by: Labbi, Safwan, et al.
Published: (2024) -
Refined Analysis of Entropy-Regularized Actor-Critic
by: Labbi, Safwan, et al.
Published: (2026) -
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
by: Mangold, Paul, et al.
Published: (2024) -
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025)