Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xingtu, Yang, Lin F., Vaswani, Sharan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026)
by: Lu, Michael, et al.
Published: (2026)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Near-Optimal Sample Complexity for Online Constrained MDPs
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025)
by: Sahu, Sharan
Published: (2025)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
by: Liu, Xingtu
Published: (2021)
by: Liu, Xingtu
Published: (2021)
An Information-Theoretic Analysis of OOD Generalization in Meta-Reinforcement Learning
by: Liu, Xingtu
Published: (2025)
by: Liu, Xingtu
Published: (2025)
Central Limit Theorems for Asynchronous Averaged Q-Learning
by: Liu, Xingtu
Published: (2025)
by: Liu, Xingtu
Published: (2025)
Sample Complexity Characterization for Linear Contextual MDPs
by: Deng, Junze, et al.
Published: (2024)
by: Deng, Junze, et al.
Published: (2024)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025)
by: Vaswani, Sharan, et al.
Published: (2025)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
by: Asad, Reza, et al.
Published: (2025)
by: Asad, Reza, et al.
Published: (2025)
From Inverse Optimization to Feasibility to ERM
by: Mishra, Saurabh, et al.
Published: (2024)
by: Mishra, Saurabh, et al.
Published: (2024)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
by: Vaswani, Sharan, et al.
Published: (2026)
by: Vaswani, Sharan, et al.
Published: (2026)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021)
by: Vaswani, Sharan, et al.
Published: (2021)
Preserving Plasticity in Continual Learning with Adaptive Linearity Injection
by: Rohani, Seyed Roozbeh Razavi, et al.
Published: (2025)
by: Rohani, Seyed Roozbeh Razavi, et al.
Published: (2025)
Towards Parameter-Free Temporal Difference Learning
by: Li, Yunxiang, et al.
Published: (2026)
by: Li, Yunxiang, et al.
Published: (2026)
Improving OOD Generalization of Pre-trained Encoders via Aligned Embedding-Space Ensembles
by: Peng, Shuman, et al.
Published: (2024)
by: Peng, Shuman, et al.
Published: (2024)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual Bandits
by: Lin, Jiabin, et al.
Published: (2024)
by: Lin, Jiabin, et al.
Published: (2024)
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
by: Fox, Curtis, et al.
Published: (2025)
by: Fox, Curtis, et al.
Published: (2025)
Time-Constrained Robust MDPs
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
by: Tan, Kevin, et al.
Published: (2024)
by: Tan, Kevin, et al.
Published: (2024)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
by: Tarbouriech, Jean, et al.
Published: (2026)
by: Tarbouriech, Jean, et al.
Published: (2026)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
by: Shehper, Ali, et al.
Published: (2026)
by: Shehper, Ali, et al.
Published: (2026)
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025)
by: Gangrade, Aditya, et al.
Published: (2025)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Similar Items
-
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026) -
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025) -
Near-Optimal Sample Complexity for Online Constrained MDPs
by: Liu, Chang, et al.
Published: (2026) -
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025) -
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)