Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kaya, Ege C., Pourghani, Aliasghar, Gupta, Vijay, Hashemi, Abolfazl |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Localized Distributional Robustness in Submodular Multi-Task Subset Selection
by: Kaya, Ege C., et al.
Published: (2024)
by: Kaya, Ege C., et al.
Published: (2024)
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026)
by: Fazla, Arda, et al.
Published: (2026)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
FedSGM: A Unified Framework for Constraint Aware, Bidirectionally Compressed, Multi-Step Federated Optimization
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
Is Bellman Equation Enough for Learning Control?
by: You, Haoxiang, et al.
Published: (2025)
by: You, Haoxiang, et al.
Published: (2025)
Randomized Greedy Methods for Weak Submodular Sensor Selection with Robustness Considerations
by: Kaya, Ege C., et al.
Published: (2024)
by: Kaya, Ege C., et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
by: Wan, Yi, et al.
Published: (2024)
by: Wan, Yi, et al.
Published: (2024)
Sampling-based Safe Reinforcement Learning for Nonlinear Dynamical Systems
by: Suttle, Wesley A., et al.
Published: (2024)
by: Suttle, Wesley A., et al.
Published: (2024)
RAMPAGE: RAndomized Mid-Point for debiAsed Gradient Extrapolation
by: Luo, Zhankun, et al.
Published: (2026)
by: Luo, Zhankun, et al.
Published: (2026)
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
by: Zhou, Angela
Published: (2024)
by: Zhou, Angela
Published: (2024)
Model-Free Learning for the Linear Quadratic Regulator over Rate-Limited Channels
by: Ye, Lintao, et al.
Published: (2024)
by: Ye, Lintao, et al.
Published: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Learning Decentralized Linear Quadratic Regulators with $\sqrt{T}$ Regret
by: Ye, Lintao, et al.
Published: (2022)
by: Ye, Lintao, et al.
Published: (2022)
Foundations of Multivariate Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
On the Foundation of Distributionally Robust Reinforcement Learning
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Tail Distribution of Regret in Optimistic Reinforcement Learning
by: Khodadadian, Sajad, et al.
Published: (2025)
by: Khodadadian, Sajad, et al.
Published: (2025)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
by: Yu, Kihyun, et al.
Published: (2024)
by: Yu, Kihyun, et al.
Published: (2024)
Submodular Maximization Approaches for Equitable Client Selection in Federated Learning
by: Jiménez, Andrés Catalino Castillo, et al.
Published: (2024)
by: Jiménez, Andrés Catalino Castillo, et al.
Published: (2024)
Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting
by: Ko, Hojin, et al.
Published: (2026)
by: Ko, Hojin, et al.
Published: (2026)
On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
by: Nanda, Phalguni, et al.
Published: (2026)
by: Nanda, Phalguni, et al.
Published: (2026)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Central Limit Theorems for Asynchronous Averaged Q-Learning
by: Liu, Xingtu
Published: (2025)
by: Liu, Xingtu
Published: (2025)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
End-to-End Learning Framework for Solving Non-Markovian Optimal Control
by: Zhang, Xiaole, et al.
Published: (2025)
by: Zhang, Xiaole, et al.
Published: (2025)
Verifiable Error Bounds for Physics-Informed Neural Network Solutions of Lyapunov and Hamilton-Jacobi-Bellman Equations
by: Liu, Jun
Published: (2026)
by: Liu, Jun
Published: (2026)
Similar Items
-
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026) -
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
by: Kaya, Ege C., et al.
Published: (2026) -
Localized Distributional Robustness in Submodular Multi-Task Subset Selection
by: Kaya, Ege C., et al.
Published: (2024) -
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026) -
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)