Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Zurek, Matthew, Chen, Yudong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Span-Based Optimal Sample Complexity for Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2023)
por: Zurek, Matthew, et al.
Publicado: (2023)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2024)
por: Zurek, Matthew, et al.
Publicado: (2024)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
por: Zurek, Matthew, et al.
Publicado: (2025)
por: Zurek, Matthew, et al.
Publicado: (2025)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
por: Zurek, Matthew, et al.
Publicado: (2024)
por: Zurek, Matthew, et al.
Publicado: (2024)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
por: Zamir, Guy, et al.
Publicado: (2026)
por: Zamir, Guy, et al.
Publicado: (2026)
Gap-Free Clustering: Sensitivity and Robustness of SDP
por: Zurek, Matthew, et al.
Publicado: (2023)
por: Zurek, Matthew, et al.
Publicado: (2023)
Faster Fixed-Point Methods for Multichain MDPs
por: Zurek, Matthew, et al.
Publicado: (2025)
por: Zurek, Matthew, et al.
Publicado: (2025)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
por: Li, Gen, et al.
Publicado: (2021)
por: Li, Gen, et al.
Publicado: (2021)
Optimal Sample Complexity for Average Reward Markov Decision Processes
por: Wang, Shengbo, et al.
Publicado: (2023)
por: Wang, Shengbo, et al.
Publicado: (2023)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
por: Chen, Zijun, et al.
Publicado: (2025)
por: Chen, Zijun, et al.
Publicado: (2025)
Stochastic Zeroth-Order Optimization under Strongly Convexity and Lipschitz Hessian: Minimax Sample Complexity
por: Yu, Qian, et al.
Publicado: (2024)
por: Yu, Qian, et al.
Publicado: (2024)
Geometry, Computation, and Optimality in Stochastic Optimization
por: Cheng, Chen, et al.
Publicado: (2019)
por: Cheng, Chen, et al.
Publicado: (2019)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
por: Chae, Woojin, et al.
Publicado: (2024)
por: Chae, Woojin, et al.
Publicado: (2024)
On the Robustness of Cross-Concentrated Sampling for Matrix Completion
por: Cai, HanQin, et al.
Publicado: (2024)
por: Cai, HanQin, et al.
Publicado: (2024)
Structured Sampling for Robust Euclidean Distance Geometry
por: Kundu, Chandra, et al.
Publicado: (2024)
por: Kundu, Chandra, et al.
Publicado: (2024)
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
por: Li, Gen, et al.
Publicado: (2023)
por: Li, Gen, et al.
Publicado: (2023)
Optimal transport natural gradient for statistical manifolds with continuous sample space
por: Chen, Yifan, et al.
Publicado: (2018)
por: Chen, Yifan, et al.
Publicado: (2018)
Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control
por: Liu, Yujie, et al.
Publicado: (2025)
por: Liu, Yujie, et al.
Publicado: (2025)
Mixing Time of the Proximal Sampler in Relative Fisher Information via Strong Data Processing Inequality
por: Wibisono, Andre
Publicado: (2025)
por: Wibisono, Andre
Publicado: (2025)
Wasserstein Distributionally Robust Estimation in High Dimensions: Performance Analysis and Optimal Hyperparameter Tuning
por: Aolaritei, Liviu, et al.
Publicado: (2022)
por: Aolaritei, Liviu, et al.
Publicado: (2022)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
por: Li, Gen, et al.
Publicado: (2020)
por: Li, Gen, et al.
Publicado: (2020)
Proximal Oracles for Optimization and Sampling
por: Liang, Jiaming, et al.
Publicado: (2024)
por: Liang, Jiaming, et al.
Publicado: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
por: Boone, Victor, et al.
Publicado: (2024)
por: Boone, Victor, et al.
Publicado: (2024)
The Oracle Complexity of Simplex-based Matrix Games
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
por: Wang, Shengbo, et al.
Publicado: (2026)
por: Wang, Shengbo, et al.
Publicado: (2026)
Fast Computation of Optimal Transport via Entropy-Regularized Extragradient Methods
por: Li, Gen, et al.
Publicado: (2023)
por: Li, Gen, et al.
Publicado: (2023)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
por: Wang, Shengbo, et al.
Publicado: (2025)
por: Wang, Shengbo, et al.
Publicado: (2025)
Optimal Online Bookmaking for Binary Games
por: Bhatt, Alankrita, et al.
Publicado: (2025)
por: Bhatt, Alankrita, et al.
Publicado: (2025)
Linear regression with overparameterized linear neural networks: Tight upper and lower bounds for implicit $\ell^1$-regularization
por: Matt, Hannes, et al.
Publicado: (2025)
por: Matt, Hannes, et al.
Publicado: (2025)
A Neural Network Algorithm for KL Divergence Estimation with Quantitative Error Bounds
por: Foss, Mikil, et al.
Publicado: (2025)
por: Foss, Mikil, et al.
Publicado: (2025)
A Dual Basis Approach for Structured Robust Euclidean Distance Geometry
por: Kundu, Chandra, et al.
Publicado: (2025)
por: Kundu, Chandra, et al.
Publicado: (2025)
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
por: Shen, Wei, et al.
Publicado: (2025)
por: Shen, Wei, et al.
Publicado: (2025)
On the Convergence Analysis of Muon
por: Shen, Wei, et al.
Publicado: (2025)
por: Shen, Wei, et al.
Publicado: (2025)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
por: Shen, Wei, et al.
Publicado: (2025)
por: Shen, Wei, et al.
Publicado: (2025)
Convexity in Disguise: A Theoretical Framework for Nonconvex Low-Rank Matrix Estimation
por: Cui, Chengyu, et al.
Publicado: (2026)
por: Cui, Chengyu, et al.
Publicado: (2026)
Recovering Simultaneously Structured Data via Non-Convex Iteratively Reweighted Least Squares
por: Kümmerle, Christian, et al.
Publicado: (2023)
por: Kümmerle, Christian, et al.
Publicado: (2023)
Generalized Orthogonal Procrustes Problem under Arbitrary Adversaries
por: Ling, Shuyang
Publicado: (2021)
por: Ling, Shuyang
Publicado: (2021)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
por: Shen, Wei, et al.
Publicado: (2023)
por: Shen, Wei, et al.
Publicado: (2023)
Tight Regret Bounds for Bayesian Optimization in One Dimension
por: Scarlett, Jonathan
Publicado: (2018)
por: Scarlett, Jonathan
Publicado: (2018)
Variational Inference on the Boolean Hypercube with the Quantum Entropy
por: Beyler, Eliot, et al.
Publicado: (2024)
por: Beyler, Eliot, et al.
Publicado: (2024)
Ejemplares similares
-
Span-Based Optimal Sample Complexity for Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2023) -
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2024) -
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
por: Zurek, Matthew, et al.
Publicado: (2025) -
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
por: Zurek, Matthew, et al.
Publicado: (2024) -
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
por: Zamir, Guy, et al.
Publicado: (2026)