Nash Learning from Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Munos, Rémi, Valko, Michal, Calandriello, Daniele, Azar, Mohammad Gheshlaghi, Rowland, Mark, Guo, Zhaohan Daniel, Tang, Yunhao, Geist, Matthieu, Mesnard, Thomas, Michi, Andrea, Selvi, Marco, Girgin, Sertan, Momchev, Nikola, Bachem, Olivier, Mankowitz, Daniel J., Precup, Doina, Piot, Bilal |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
VA-learning as a more efficient alternative to Q-learning
by: Tang, Yunhao, et al.
Published: (2023)
by: Tang, Yunhao, et al.
Published: (2023)
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024)
by: Calandriello, Daniele, et al.
Published: (2024)
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023)
by: Rowland, Mark, et al.
Published: (2023)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Offline Regularised Reinforcement Learning for Large Language Models Alignment
by: Richemond, Pierre Harvey, et al.
Published: (2024)
by: Richemond, Pierre Harvey, et al.
Published: (2024)
Bandits attack function optimization
by: Preux, Philippe, et al.
Published: (2026)
by: Preux, Philippe, et al.
Published: (2026)
Stochastic simultaneous optimistic optimization
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Black-box optimization of noisy functions with unknown smoothness
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Spectral Thompson sampling
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
by: Rowland, Mark, et al.
Published: (2024)
by: Rowland, Mark, et al.
Published: (2024)
Model-free Posterior Sampling via Learning Rate Randomization
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Diversity-Rewarded CFG Distillation
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
Large-scale semi-supervised learning with online spectral graph sparsification
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Analysis of Nystrom method with sequential ridge leverage scores
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Diversity-Enriched Option-Critic
by: Kamat, Anand, et al.
Published: (2020)
by: Kamat, Anand, et al.
Published: (2020)
Functional Acceleration for Policy Mirror Descent
by: Chelu, Veronica, et al.
Published: (2024)
by: Chelu, Veronica, et al.
Published: (2024)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
by: Alver, Safa, et al.
Published: (2022)
by: Alver, Safa, et al.
Published: (2022)
Proximal Point Nash Learning from Human Feedback
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Optimizing Return Distributions with Distributional Dynamic Programming
by: Pires, Bernardo Ávila, et al.
Published: (2025)
by: Pires, Bernardo Ávila, et al.
Published: (2025)
Spectral bandits for smooth graph functions with applications in recommender systems
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Learning in Mean Field Games: A Survey
by: Laurière, Mathieu, et al.
Published: (2022)
by: Laurière, Mathieu, et al.
Published: (2022)
Multi-turn Reinforcement Learning from Preference Human Feedback
by: Shani, Lior, et al.
Published: (2024)
by: Shani, Lior, et al.
Published: (2024)
Improved large-scale graph learning through ridge spectral sparsification
by: Calandriello, Daniele, et al.
Published: (2026)
by: Calandriello, Daniele, et al.
Published: (2026)
Planning in entropy-regularized Markov decision processes and games
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Balancing Plasticity and Stability with Fast and Slow Successor Features
by: Chua, Raymond, et al.
Published: (2026)
by: Chua, Raymond, et al.
Published: (2026)
On the Privacy of Selection Mechanisms with Gaussian Noise
by: Lebensold, Jonathan, et al.
Published: (2024)
by: Lebensold, Jonathan, et al.
Published: (2024)
MusicRL: Aligning Music Generation to Human Preferences
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
Similar Items
-
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024) -
VA-learning as a more efficient alternative to Q-learning
by: Tang, Yunhao, et al.
Published: (2023) -
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024) -
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023) -
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)