Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Tian, Yang, Lin F., Szepesvári, Csaba |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024)
by: Tkachuk, Volodymyr, et al.
Published: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Rectifying Regression in Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Sharp analysis of linear ensemble sampling
by: Akhavan, Arya, et al.
Published: (2026)
by: Akhavan, Arya, et al.
Published: (2026)
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
Near-Optimal Sample Complexity for Online Constrained MDPs
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Convergence of Natural Policy Gradient for a Family of Infinite-State Queueing MDPs
by: Grosof, Isaac, et al.
Published: (2024)
by: Grosof, Isaac, et al.
Published: (2024)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
Balancing optimism and pessimism in offline-to-online learning
by: Sentenac, Flore, et al.
Published: (2025)
by: Sentenac, Flore, et al.
Published: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
by: Liu, Xingtu, et al.
Published: (2025)
by: Liu, Xingtu, et al.
Published: (2025)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Exploration via linearly perturbed loss minimisation
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Stochastic Gradient Descent for Gaussian Processes Done Right
by: Lin, Jihao Andreas, et al.
Published: (2023)
by: Lin, Jihao Andreas, et al.
Published: (2023)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Time-Constrained Robust MDPs
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026)
by: Bakhtiari, Alireza, et al.
Published: (2026)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026)
by: Lu, Michael, et al.
Published: (2026)
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Regret Minimization via Saddle Point Optimization
by: Kirschner, Johannes, et al.
Published: (2024)
by: Kirschner, Johannes, et al.
Published: (2024)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Similar Items
-
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021) -
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024) -
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025) -
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
by: Liu, Chang, et al.
Published: (2026) -
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)