VA-learning as a more efficient alternative to Q-learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Yunhao, Munos, Rémi, Rowland, Mark, Valko, Michal |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Bandits attack function optimization
by: Preux, Philippe, et al.
Published: (2026)
by: Preux, Philippe, et al.
Published: (2026)
Stochastic simultaneous optimistic optimization
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Black-box optimization of noisy functions with unknown smoothness
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
by: Rowland, Mark, et al.
Published: (2024)
by: Rowland, Mark, et al.
Published: (2024)
Spectral Thompson sampling
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Distance metric learning for conditional anomaly detection
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Planning in entropy-regularized Markov decision processes and games
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Online learning with noisy side observations
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Online learning with Erdős-Rényi side-observation graphs
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Adaptive multi-fidelity optimization with fast learning rates
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023)
by: Rowland, Mark, et al.
Published: (2023)
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024)
by: Calandriello, Daniele, et al.
Published: (2024)
Large-scale semi-supervised learning with online spectral graph sparsification
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Semi-supervised learning with max-margin graph cuts
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
On two ways to use determinantal point processes for Monte Carlo integration
by: Gautier, Guillaume, et al.
Published: (2026)
by: Gautier, Guillaume, et al.
Published: (2026)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Improved large-scale graph learning through ridge spectral sparsification
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Bandits on graphs and structures
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Online semi-supervised perception: Real-time learning without explicit feedback
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Middle-mile logistics through the lens of goal-conditioned reinforcement learning
by: Eberhard, Onno, et al.
Published: (2026)
by: Eberhard, Onno, et al.
Published: (2026)
Model-free Posterior Sampling via Learning Rate Randomization
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Super-Exponential Regret for UCT, AlphaGo and Variants
by: Orseau, Laurent, et al.
Published: (2024)
by: Orseau, Laurent, et al.
Published: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Learning from a single labeled face and a stream of unlabeled data
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Feature importance analysis for patient management decisions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
by: Neu, Gergely, et al.
Published: (2026)
by: Neu, Gergely, et al.
Published: (2026)
Similar Items
-
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024) -
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026) -
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
by: Grill, Jean-Bastien, et al.
Published: (2026) -
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025) -
Bandits attack function optimization
by: Preux, Philippe, et al.
Published: (2026)