Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Max Qiushi, Asad, Reza, Tan, Kevin, Ishfaq, Haque, Szepesvari, Csaba, Vaswani, Sharan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025)
von: Asad, Reza, et al.
Veröffentlicht: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
von: Zhong, Han, et al.
Veröffentlicht: (2021)
von: Zhong, Han, et al.
Veröffentlicht: (2021)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
von: Lu, Michael, et al.
Veröffentlicht: (2026)
von: Lu, Michael, et al.
Veröffentlicht: (2026)
Fast Convergence of Softmax Policy Mirror Ascent
von: Asad, Reza, et al.
Veröffentlicht: (2024)
von: Asad, Reza, et al.
Veröffentlicht: (2024)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
von: Ishfaq, Haque, et al.
Veröffentlicht: (2025)
von: Ishfaq, Haque, et al.
Veröffentlicht: (2025)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
von: Tian, Tian, et al.
Veröffentlicht: (2024)
von: Tian, Tian, et al.
Veröffentlicht: (2024)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
von: Dang, Anh, et al.
Veröffentlicht: (2024)
von: Dang, Anh, et al.
Veröffentlicht: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
Towards Parameter-Free Temporal Difference Learning
von: Li, Yunxiang, et al.
Veröffentlicht: (2026)
von: Li, Yunxiang, et al.
Veröffentlicht: (2026)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
von: Lu, Michael, et al.
Veröffentlicht: (2024)
von: Lu, Michael, et al.
Veröffentlicht: (2024)
Optimistic Regret Bounds for Online Learning in Adversarial Markov Decision Processes
von: Moon, Sang Bin, et al.
Veröffentlicht: (2024)
von: Moon, Sang Bin, et al.
Veröffentlicht: (2024)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
von: Sherman, Uri, et al.
Veröffentlicht: (2023)
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Sharp analysis of linear ensemble sampling
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes
von: Kone, Cyrille, et al.
Veröffentlicht: (2026)
von: Kone, Cyrille, et al.
Veröffentlicht: (2026)
Generalized Linear Markov Decision Process
von: Zhang, Sinian, et al.
Veröffentlicht: (2025)
von: Zhang, Sinian, et al.
Veröffentlicht: (2025)
Balancing optimism and pessimism in offline-to-online learning
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
From Inverse Optimization to Feasibility to ERM
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
von: Zeng, Sihan, et al.
Veröffentlicht: (2021)
von: Zeng, Sihan, et al.
Veröffentlicht: (2021)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation
von: Gu, Jingwen, et al.
Veröffentlicht: (2025)
von: Gu, Jingwen, et al.
Veröffentlicht: (2025)
Policy Testing in Markov Decision Processes
von: Ariu, Kaito, et al.
Veröffentlicht: (2025)
von: Ariu, Kaito, et al.
Veröffentlicht: (2025)
Exploration via linearly perturbed loss minimisation
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Optimal Decision Tree Policies for Markov Decision Processes
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
Preserving Plasticity in Continual Learning with Adaptive Linearity Injection
von: Rohani, Seyed Roozbeh Razavi, et al.
Veröffentlicht: (2025)
von: Rohani, Seyed Roozbeh Razavi, et al.
Veröffentlicht: (2025)
Horizon-Free Regret for Linear Markov Decision Processes
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
Achieving Constant Regret in Linear Markov Decision Processes
von: Zhang, Weitong, et al.
Veröffentlicht: (2024)
von: Zhang, Weitong, et al.
Veröffentlicht: (2024)
Actor-Critics Can Achieve Optimal Sample Efficiency
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025) -
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025) -
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
von: Zhong, Han, et al.
Veröffentlicht: (2021) -
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
von: Lu, Michael, et al.
Veröffentlicht: (2026) -
Fast Convergence of Softmax Policy Mirror Ascent
von: Asad, Reza, et al.
Veröffentlicht: (2024)