Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Boone, Victor, Zhang, Zihan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Distributionally Robust Regret Optimal Control Under Moment-Based Ambiguity Sets
by: Taha, Feras Al, et al.
Published: (2025)
by: Taha, Feras Al, et al.
Published: (2025)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Model approximation in MDPs with unbounded per-step cost
by: Bozkurt, Berk, et al.
Published: (2024)
by: Bozkurt, Berk, et al.
Published: (2024)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Scalable spectral representations for multi-agent reinforcement learning in network MDPs
by: Ren, Zhaolin, et al.
Published: (2024)
by: Ren, Zhaolin, et al.
Published: (2024)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Regret Analysis: a control perspective
by: Gibson, Travis E., et al.
Published: (2025)
by: Gibson, Travis E., et al.
Published: (2025)
Almost Surely $\sqrt{T}$ Regret for Adaptive LQR
by: Lu, Yiwen, et al.
Published: (2023)
by: Lu, Yiwen, et al.
Published: (2023)
Learning Decentralized Linear Quadratic Regulators with $\sqrt{T}$ Regret
by: Ye, Lintao, et al.
Published: (2022)
by: Ye, Lintao, et al.
Published: (2022)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
by: Zamir, Guy, et al.
Published: (2026)
by: Zamir, Guy, et al.
Published: (2026)
Regret Analysis of Policy Optimization over Submanifolds for Linearly Constrained Online LQG
by: Chang, Ting-Jui, et al.
Published: (2024)
by: Chang, Ting-Jui, et al.
Published: (2024)
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
by: Sattar, Yahya, et al.
Published: (2021)
by: Sattar, Yahya, et al.
Published: (2021)
Online Nonstochastic Prediction: Logarithmic Regret via Predictive Online Least Squares
by: Pai, Chih-Fan, et al.
Published: (2026)
by: Pai, Chih-Fan, et al.
Published: (2026)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
Robust Q-Learning under Corrupted Rewards
by: Maity, Sreejeet, et al.
Published: (2024)
by: Maity, Sreejeet, et al.
Published: (2024)
Optimistic Online LQR via Intrinsic Rewards
by: Bartos, Marcell, et al.
Published: (2026)
by: Bartos, Marcell, et al.
Published: (2026)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Hybrid Energy-Aware Reward Shaping: A Unified Lightweight Physics-Guided Methodology for Policy Optimization
by: Liao, Qijun, et al.
Published: (2026)
by: Liao, Qijun, et al.
Published: (2026)
End-to-End Learning Framework for Solving Non-Markovian Optimal Control
by: Zhang, Xiaole, et al.
Published: (2025)
by: Zhang, Xiaole, et al.
Published: (2025)
Approximate Solution Methods for the Average Reward Criterion in Optimal Tracking Control of Linear Systems
by: Nguyen, Duc Cuong
Published: (2025)
by: Nguyen, Duc Cuong
Published: (2025)
Robust Regret Optimal Control
by: Liu, Jietian, et al.
Published: (2023)
by: Liu, Jietian, et al.
Published: (2023)
The Ground Cost for Optimal Transport of Angular Velocity
by: Elamvazhuthi, Karthik, et al.
Published: (2025)
by: Elamvazhuthi, Karthik, et al.
Published: (2025)
Deterministic Trajectory Optimization through Probabilistic Optimal Control
by: Filabadi, Mohammad Mahmoudi, et al.
Published: (2024)
by: Filabadi, Mohammad Mahmoudi, et al.
Published: (2024)
Towards an Optimal Control Perspective of ResNet Training
by: Püttschneider, Jens, et al.
Published: (2025)
by: Püttschneider, Jens, et al.
Published: (2025)
Optimal Centered Active Excitation in Linear System Identification
by: Ito, Kaito, et al.
Published: (2026)
by: Ito, Kaito, et al.
Published: (2026)
A Minimax Optimal Controller for Positive Systems
by: Gurpegui, Alba, et al.
Published: (2025)
by: Gurpegui, Alba, et al.
Published: (2025)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Mutual Information Optimal Control of Discrete-Time Linear Systems
by: Enami, Shoju, et al.
Published: (2025)
by: Enami, Shoju, et al.
Published: (2025)
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
by: Maity, Sreejeet, et al.
Published: (2025)
by: Maity, Sreejeet, et al.
Published: (2025)
On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
by: Enami, Shoju, et al.
Published: (2025)
by: Enami, Shoju, et al.
Published: (2025)
InterQ: A DQN Framework for Optimal Intermittent Control
by: Aggarwal, Shubham, et al.
Published: (2025)
by: Aggarwal, Shubham, et al.
Published: (2025)
CLT-Optimal Parameter Error Bounds for Linear System Identification
by: Zhou, Yichen, et al.
Published: (2026)
by: Zhou, Yichen, et al.
Published: (2026)
Similar Items
-
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025) -
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026) -
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023) -
Distributionally Robust Regret Optimal Control Under Moment-Based Ambiguity Sets
by: Taha, Feras Al, et al.
Published: (2025) -
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)