Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
Fuente:
arXiv
Salvato in:
| Autori principali: | Labbi, Safwan, Tiapkin, Daniil, Mangold, Paul, Moulines, Eric |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
di: Labbi, Safwan, et al.
Pubblicazione: (2025)
di: Labbi, Safwan, et al.
Pubblicazione: (2025)
Refined Analysis of Entropy-Regularized Actor-Critic
di: Labbi, Safwan, et al.
Pubblicazione: (2026)
di: Labbi, Safwan, et al.
Pubblicazione: (2026)
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
di: Labbi, Safwan, et al.
Pubblicazione: (2024)
di: Labbi, Safwan, et al.
Pubblicazione: (2024)
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
di: Ocello, Antonio, et al.
Pubblicazione: (2025)
di: Ocello, Antonio, et al.
Pubblicazione: (2025)
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
di: Mangold, Paul, et al.
Pubblicazione: (2024)
di: Mangold, Paul, et al.
Pubblicazione: (2024)
Joint Channel Selection using FedDRL in V2X
di: Mancini, Lorenzo, et al.
Pubblicazione: (2024)
di: Mancini, Lorenzo, et al.
Pubblicazione: (2024)
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
di: Mangold, Paul, et al.
Pubblicazione: (2025)
di: Mangold, Paul, et al.
Pubblicazione: (2025)
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
di: Ding, Yuhao, et al.
Pubblicazione: (2021)
di: Ding, Yuhao, et al.
Pubblicazione: (2021)
Generative Flow Networks as Entropy-Regularized RL
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
Demonstration-Regularized RL
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
di: Samsonov, Sergey, et al.
Pubblicazione: (2023)
di: Samsonov, Sergey, et al.
Pubblicazione: (2023)
Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up
di: Mangold, Paul, et al.
Pubblicazione: (2025)
di: Mangold, Paul, et al.
Pubblicazione: (2025)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
di: Indelman, Hedda Cohen, et al.
Pubblicazione: (2024)
di: Indelman, Hedda Cohen, et al.
Pubblicazione: (2024)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
di: Klein, Sara, et al.
Pubblicazione: (2023)
di: Klein, Sara, et al.
Pubblicazione: (2023)
Model-free Posterior Sampling via Learning Rate Randomization
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2023)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
di: Tiapkin, Daniil, et al.
Pubblicazione: (2024)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
di: Lin, Max Qiushi, et al.
Pubblicazione: (2025)
di: Lin, Max Qiushi, et al.
Pubblicazione: (2025)
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
di: Zhu, Zhaoyu, et al.
Pubblicazione: (2026)
di: Zhu, Zhaoyu, et al.
Pubblicazione: (2026)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
di: Gritsaev, Timofei, et al.
Pubblicazione: (2024)
di: Gritsaev, Timofei, et al.
Pubblicazione: (2024)
Proximal Point Nash Learning from Human Feedback
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
di: Manenti, Alessandro, et al.
Pubblicazione: (2025)
di: Manenti, Alessandro, et al.
Pubblicazione: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
di: Levin, Ilya, et al.
Pubblicazione: (2026)
di: Levin, Ilya, et al.
Pubblicazione: (2026)
Refined Analysis of Federated Averaging and Federated Richardson-Romberg
di: Mangold, Paul, et al.
Pubblicazione: (2024)
di: Mangold, Paul, et al.
Pubblicazione: (2024)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
di: Cayci, Semih, et al.
Pubblicazione: (2021)
di: Cayci, Semih, et al.
Pubblicazione: (2021)
Linear Convergence of Independent Natural Policy Gradient in Games with Entropy Regularization
di: Sun, Youbang, et al.
Pubblicazione: (2024)
di: Sun, Youbang, et al.
Pubblicazione: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
di: Asad, Reza, et al.
Pubblicazione: (2024)
di: Asad, Reza, et al.
Pubblicazione: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
di: Li, Yingru
Pubblicazione: (2025)
di: Li, Yingru
Pubblicazione: (2025)
Revisiting Non-Acyclic GFlowNets in Discrete Environments
di: Morozov, Nikita, et al.
Pubblicazione: (2025)
di: Morozov, Nikita, et al.
Pubblicazione: (2025)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
di: Ged, François, et al.
Pubblicazione: (2023)
di: Ged, François, et al.
Pubblicazione: (2023)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
di: Lattimore, Tor
Pubblicazione: (2026)
di: Lattimore, Tor
Pubblicazione: (2026)
Beyond Softmax: A New Perspective on Gradient Bandits
di: Melo, Emerson, et al.
Pubblicazione: (2025)
di: Melo, Emerson, et al.
Pubblicazione: (2025)
Incentivized Learning in Principal-Agent Bandit Games
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
di: Scheid, Antoine, et al.
Pubblicazione: (2024)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
di: Varre, Aditya, et al.
Pubblicazione: (2026)
di: Varre, Aditya, et al.
Pubblicazione: (2026)
Learning Shortest Paths with Generative Flow Networks
di: Morozov, Nikita, et al.
Pubblicazione: (2026)
di: Morozov, Nikita, et al.
Pubblicazione: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
Adaptive Set-Mass Calibration with Conformal Prediction
di: Kazantsev, Daniil, et al.
Pubblicazione: (2025)
di: Kazantsev, Daniil, et al.
Pubblicazione: (2025)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
di: Kar, Avik, et al.
Pubblicazione: (2026)
di: Kar, Avik, et al.
Pubblicazione: (2026)
Accelerated Policy Gradient: On the Convergence Rates of the Nesterov Momentum for Reinforcement Learning
di: Chen, Yen-Ju, et al.
Pubblicazione: (2023)
di: Chen, Yen-Ju, et al.
Pubblicazione: (2023)
Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures
di: Zhou, Mo, et al.
Pubblicazione: (2025)
di: Zhou, Mo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
di: Labbi, Safwan, et al.
Pubblicazione: (2025) -
Refined Analysis of Entropy-Regularized Actor-Critic
di: Labbi, Safwan, et al.
Pubblicazione: (2026) -
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
di: Labbi, Safwan, et al.
Pubblicazione: (2024) -
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
di: Ocello, Antonio, et al.
Pubblicazione: (2025) -
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
di: Mangold, Paul, et al.
Pubblicazione: (2024)