Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
Fuente:
arXiv
Saved in:
| Main Authors: | Rowland, Mark, Wenliang, Li Kevin, Munos, Rémi, Lyle, Clare, Tang, Yunhao, Dabney, Will |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
VA-learning as a more efficient alternative to Q-learning
by: Tang, Yunhao, et al.
Published: (2023)
by: Tang, Yunhao, et al.
Published: (2023)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023)
by: Rowland, Mark, et al.
Published: (2023)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning
by: Khetarpal, Khimya, et al.
Published: (2024)
by: Khetarpal, Khimya, et al.
Published: (2024)
Optimizing Return Distributions with Distributional Dynamic Programming
by: Pires, Bernardo Ávila, et al.
Published: (2025)
by: Pires, Bernardo Ávila, et al.
Published: (2025)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
A Distributional Analogue to the Successor Representation
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
by: Lee, Harin, et al.
Published: (2026)
by: Lee, Harin, et al.
Published: (2026)
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
by: Tajdini, Artin, et al.
Published: (2023)
by: Tajdini, Artin, et al.
Published: (2023)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Minimax-Optimal Multi-Agent Robust Reinforcement Learning
by: Jiao, Yuchen, et al.
Published: (2024)
by: Jiao, Yuchen, et al.
Published: (2024)
Distributional Bellman Operators over Mean Embeddings
by: Wenliang, Li Kevin, et al.
Published: (2023)
by: Wenliang, Li Kevin, et al.
Published: (2023)
Foundations of Multivariate Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
by: Beukman, Michael, et al.
Published: (2026)
by: Beukman, Michael, et al.
Published: (2026)
Super-Exponential Regret for UCT, AlphaGo and Variants
by: Orseau, Laurent, et al.
Published: (2024)
by: Orseau, Laurent, et al.
Published: (2024)
Near-Optimal Distributed Minimax Optimization under the Second-Order Similarity
by: Zhou, Qihao, et al.
Published: (2024)
by: Zhou, Qihao, et al.
Published: (2024)
Lifelong Reinforcement Learning via Neuromodulation
by: Lee, Sebastian, et al.
Published: (2024)
by: Lee, Sebastian, et al.
Published: (2024)
Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning
by: Liu, Zhishuai, et al.
Published: (2024)
by: Liu, Zhishuai, et al.
Published: (2024)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
by: Huang, Yan, et al.
Published: (2024)
by: Huang, Yan, et al.
Published: (2024)
Bandits attack function optimization
by: Preux, Philippe, et al.
Published: (2026)
by: Preux, Philippe, et al.
Published: (2026)
Stochastic simultaneous optimistic optimization
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
by: Ghosh, Debamita, et al.
Published: (2025)
by: Ghosh, Debamita, et al.
Published: (2025)
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024)
by: Calandriello, Daniele, et al.
Published: (2024)
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity
by: Shi, Laixi, et al.
Published: (2022)
by: Shi, Laixi, et al.
Published: (2022)
Disentangling the Causes of Plasticity Loss in Neural Networks
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Black-box optimization of noisy functions with unknown smoothness
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Outcome-based Exploration for LLM Reasoning
by: Song, Yuda, et al.
Published: (2025)
by: Song, Yuda, et al.
Published: (2025)
Similar Items
-
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024) -
VA-learning as a more efficient alternative to Q-learning
by: Tang, Yunhao, et al.
Published: (2023) -
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025) -
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025) -
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023)