Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zurek, Matthew, Chen, Yudong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Span-Based Optimal Sample Complexity for Average Reward MDPs
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
di: Zamir, Guy, et al.
Pubblicazione: (2026)
di: Zamir, Guy, et al.
Pubblicazione: (2026)
Faster Fixed-Point Methods for Multichain MDPs
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
Gap-Free Clustering: Sensitivity and Robustness of SDP
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
di: Chae, Woojin, et al.
Pubblicazione: (2024)
di: Chae, Woojin, et al.
Pubblicazione: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
di: Boone, Victor, et al.
Pubblicazione: (2024)
di: Boone, Victor, et al.
Pubblicazione: (2024)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
di: Li, Gen, et al.
Pubblicazione: (2021)
di: Li, Gen, et al.
Pubblicazione: (2021)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
di: Wang, Shengbo, et al.
Pubblicazione: (2026)
di: Wang, Shengbo, et al.
Pubblicazione: (2026)
Optimal Sample Complexity for Average Reward Markov Decision Processes
di: Wang, Shengbo, et al.
Pubblicazione: (2023)
di: Wang, Shengbo, et al.
Pubblicazione: (2023)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
di: Li, Gen, et al.
Pubblicazione: (2020)
di: Li, Gen, et al.
Pubblicazione: (2020)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
di: Chen, Zijun, et al.
Pubblicazione: (2025)
di: Chen, Zijun, et al.
Pubblicazione: (2025)
Stochastic Zeroth-Order Optimization under Strongly Convexity and Lipschitz Hessian: Minimax Sample Complexity
di: Yu, Qian, et al.
Pubblicazione: (2024)
di: Yu, Qian, et al.
Pubblicazione: (2024)
Geometry, Computation, and Optimality in Stochastic Optimization
di: Cheng, Chen, et al.
Pubblicazione: (2019)
di: Cheng, Chen, et al.
Pubblicazione: (2019)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
di: Wan, Yi, et al.
Pubblicazione: (2024)
di: Wan, Yi, et al.
Pubblicazione: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
di: Yu, Kihyun, et al.
Pubblicazione: (2026)
di: Yu, Kihyun, et al.
Pubblicazione: (2026)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
di: Omidi, Saber, et al.
Pubblicazione: (2025)
di: Omidi, Saber, et al.
Pubblicazione: (2025)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
di: Zhang, Runyu, et al.
Pubblicazione: (2023)
di: Zhang, Runyu, et al.
Pubblicazione: (2023)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
On the Robustness of Cross-Concentrated Sampling for Matrix Completion
di: Cai, HanQin, et al.
Pubblicazione: (2024)
di: Cai, HanQin, et al.
Pubblicazione: (2024)
Structured Sampling for Robust Euclidean Distance Geometry
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
Optimal transport natural gradient for statistical manifolds with continuous sample space
di: Chen, Yifan, et al.
Pubblicazione: (2018)
di: Chen, Yifan, et al.
Pubblicazione: (2018)
Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control
di: Liu, Yujie, et al.
Pubblicazione: (2025)
di: Liu, Yujie, et al.
Pubblicazione: (2025)
Wasserstein Distributionally Robust Estimation in High Dimensions: Performance Analysis and Optimal Hyperparameter Tuning
di: Aolaritei, Liviu, et al.
Pubblicazione: (2022)
di: Aolaritei, Liviu, et al.
Pubblicazione: (2022)
Planning and Learning in Average Risk-aware MDPs
di: Wang, Weikai, et al.
Pubblicazione: (2025)
di: Wang, Weikai, et al.
Pubblicazione: (2025)
Generalized Orthogonal Procrustes Problem under Arbitrary Adversaries
di: Ling, Shuyang
Pubblicazione: (2021)
di: Ling, Shuyang
Pubblicazione: (2021)
Effective Communication with Dynamic Feature Compression
di: Talli, Pietro, et al.
Pubblicazione: (2024)
di: Talli, Pietro, et al.
Pubblicazione: (2024)
Fast Computation of Optimal Transport via Entropy-Regularized Extragradient Methods
di: Li, Gen, et al.
Pubblicazione: (2023)
di: Li, Gen, et al.
Pubblicazione: (2023)
Distributed Adaptive Learning Under Communication Constraints
di: Carpentiero, Marco, et al.
Pubblicazione: (2021)
di: Carpentiero, Marco, et al.
Pubblicazione: (2021)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
Optimal Online Bookmaking for Binary Games
di: Bhatt, Alankrita, et al.
Pubblicazione: (2025)
di: Bhatt, Alankrita, et al.
Pubblicazione: (2025)
Variational Inference on the Boolean Hypercube with the Quantum Entropy
di: Beyler, Eliot, et al.
Pubblicazione: (2024)
di: Beyler, Eliot, et al.
Pubblicazione: (2024)
Group Projected Subspace Pursuit for Block Sparse Signal Reconstruction: Convergence Analysis and Applications
di: He, Roy Y., et al.
Pubblicazione: (2024)
di: He, Roy Y., et al.
Pubblicazione: (2024)
Convexity in Disguise: A Theoretical Framework for Nonconvex Low-Rank Matrix Estimation
di: Cui, Chengyu, et al.
Pubblicazione: (2026)
di: Cui, Chengyu, et al.
Pubblicazione: (2026)
Linear regression with overparameterized linear neural networks: Tight upper and lower bounds for implicit $\ell^1$-regularization
di: Matt, Hannes, et al.
Pubblicazione: (2025)
di: Matt, Hannes, et al.
Pubblicazione: (2025)
Recovering Simultaneously Structured Data via Non-Convex Iteratively Reweighted Least Squares
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
A Neural Network Algorithm for KL Divergence Estimation with Quantitative Error Bounds
di: Foss, Mikil, et al.
Pubblicazione: (2025)
di: Foss, Mikil, et al.
Pubblicazione: (2025)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
di: Shen, Wei, et al.
Pubblicazione: (2023)
di: Shen, Wei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Span-Based Optimal Sample Complexity for Average Reward MDPs
di: Zurek, Matthew, et al.
Pubblicazione: (2023) -
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
di: Zurek, Matthew, et al.
Pubblicazione: (2024) -
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025) -
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025) -
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
di: Zamir, Guy, et al.
Pubblicazione: (2026)