Ordering-based Conditions for Global Convergence of Policy Gradient Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Mei, Jincheng, Dai, Bo, Agarwal, Alekh, Ghavamzadeh, Mohammad, Szepesvari, Csaba, Schuurmans, Dale |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Rectifying Regression in Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
by: Dai, Bo, et al.
Published: (2026)
by: Dai, Bo, et al.
Published: (2026)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
by: Agarwal, Alekh, et al.
Published: (2022)
by: Agarwal, Alekh, et al.
Published: (2022)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024)
by: Tkachuk, Volodymyr, et al.
Published: (2024)
Sharp analysis of linear ensemble sampling
by: Akhavan, Arya, et al.
Published: (2026)
by: Akhavan, Arya, et al.
Published: (2026)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
Balancing optimism and pessimism in offline-to-online learning
by: Sentenac, Flore, et al.
Published: (2025)
by: Sentenac, Flore, et al.
Published: (2025)
On the Global Convergence of Risk-Averse Natural Policy Gradient Methods with Expected Conditional Risk Measures
by: Yu, Xian, et al.
Published: (2023)
by: Yu, Xian, et al.
Published: (2023)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Exploration via linearly perturbed loss minimisation
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Delightful Gradients Accelerate Corner Escape
by: Mei, Jincheng, et al.
Published: (2026)
by: Mei, Jincheng, et al.
Published: (2026)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025)
by: Agarwal, Alekh, et al.
Published: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
by: Che, Fengdi, et al.
Published: (2024)
by: Che, Fengdi, et al.
Published: (2024)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
by: Hau, Jia Lin, et al.
Published: (2024)
by: Hau, Jia Lin, et al.
Published: (2024)
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026)
by: Bakhtiari, Alireza, et al.
Published: (2026)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Stochastic Gradient Descent for Gaussian Processes Done Right
by: Lin, Jihao Andreas, et al.
Published: (2023)
by: Lin, Jihao Andreas, et al.
Published: (2023)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Contextual Bandits with Stage-wise Constraints
by: Pacchiano, Aldo, et al.
Published: (2024)
by: Pacchiano, Aldo, et al.
Published: (2024)
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
by: Zhu, Zhaoyu, et al.
Published: (2026)
by: Zhu, Zhaoyu, et al.
Published: (2026)
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
Plastic Learning with Deep Fourier Features
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Similar Items
-
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024) -
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025) -
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025) -
Rectifying Regression in Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025) -
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)