RL-finetuning LLMs from on- and off-policy data with a single algorithm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yunhao, Cohen, Taco, Zhang, David W., Valko, Michal, Munos, Rémi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VA-learning as a more efficient alternative to Q-learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2023)
von: Tang, Yunhao, et al.
Veröffentlicht: (2023)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
On a few pitfalls in KL divergence gradient estimation for RL
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Efficient RL Training for LLMs with Experience Replay
von: Arnal, Charles, et al.
Veröffentlicht: (2026)
von: Arnal, Charles, et al.
Veröffentlicht: (2026)
Bandits attack function optimization
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
Stochastic simultaneous optimistic optimization
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
Black-box optimization of noisy functions with unknown smoothness
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Spectral Thompson sampling
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
Spectral bandits for smooth graph functions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Spectral bandits
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Bayesian policy gradient and actor-critic algorithms
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
Learning from a single labeled face and a stream of unlabeled data
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
Understanding the performance gap between online and offline alignment algorithms
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Planning in entropy-regularized Markov decision processes and games
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
von: Arnal, Charles, et al.
Veröffentlicht: (2025)
von: Arnal, Charles, et al.
Veröffentlicht: (2025)
A single algorithm for both restless and rested rotting bandits
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
von: Rowland, Mark, et al.
Veröffentlicht: (2024)
von: Rowland, Mark, et al.
Veröffentlicht: (2024)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
von: Sancaktar, Cansu, et al.
Veröffentlicht: (2026)
von: Sancaktar, Cansu, et al.
Veröffentlicht: (2026)
On two ways to use determinantal point processes for Monte Carlo integration
von: Gautier, Guillaume, et al.
Veröffentlicht: (2026)
von: Gautier, Guillaume, et al.
Veröffentlicht: (2026)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
Bandits on graphs and structures
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
Model-free Posterior Sampling via Learning Rate Randomization
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
Super-Exponential Regret for UCT, AlphaGo and Variants
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
Learning predictive models for combinations of heterogeneous proteomic data sources
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Feature importance analysis for patient management decisions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
Distance metric learning for conditional anomaly detection
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Revealing graph bandits for maximizing local influence
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Extreme bandits
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Human Alignment of Large Language Models through Online Preference Optimisation
von: Calandriello, Daniele, et al.
Veröffentlicht: (2024)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2024)
Demonstration-Regularized RL
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VA-learning as a more efficient alternative to Q-learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2023) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025) -
On a few pitfalls in KL divergence gradient estimation for RL
von: Tang, Yunhao, et al.
Veröffentlicht: (2025) -
Efficient RL Training for LLMs with Experience Replay
von: Arnal, Charles, et al.
Veröffentlicht: (2026) -
Bandits attack function optimization
von: Preux, Philippe, et al.
Veröffentlicht: (2026)