Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Noukhovitch, Michael, Huang, Shengyi, Xhonneux, Sophie, Hosseini, Arian, Agarwal, Rishabh, Courville, Aaron |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
V-STaR: Training Verifiers for Self-Taught Reasoners
por: Hosseini, Arian, et al.
Publicado: (2024)
por: Hosseini, Arian, et al.
Publicado: (2024)
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
por: Chandra, Abhranil, et al.
Publicado: (2025)
por: Chandra, Abhranil, et al.
Publicado: (2025)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
por: Lavoie, Samuel, et al.
Publicado: (2025)
por: Lavoie, Samuel, et al.
Publicado: (2025)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
por: Bansal, Hritik, et al.
Publicado: (2024)
por: Bansal, Hritik, et al.
Publicado: (2024)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
por: Li, Aaron J., et al.
Publicado: (2024)
por: Li, Aaron J., et al.
Publicado: (2024)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
por: Honari, Homayoun, et al.
Publicado: (2026)
por: Honari, Homayoun, et al.
Publicado: (2026)
The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
por: Huang, Shengyi, et al.
Publicado: (2024)
por: Huang, Shengyi, et al.
Publicado: (2024)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
por: Sareen, Kusha, et al.
Publicado: (2025)
por: Sareen, Kusha, et al.
Publicado: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
por: Agarwal, Rishabh, et al.
Publicado: (2023)
por: Agarwal, Rishabh, et al.
Publicado: (2023)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
por: Gao, Zhaolin, et al.
Publicado: (2024)
por: Gao, Zhaolin, et al.
Publicado: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Dataset Reset Policy Optimization for RLHF
por: Chang, Jonathan D., et al.
Publicado: (2024)
por: Chang, Jonathan D., et al.
Publicado: (2024)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
por: Huang, Jiawei, et al.
Publicado: (2025)
por: Huang, Jiawei, et al.
Publicado: (2025)
Prototypical Reward Network for Data-Efficient RLHF
por: Zhang, Jinghan, et al.
Publicado: (2024)
por: Zhang, Jinghan, et al.
Publicado: (2024)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
por: Setlur, Amrith, et al.
Publicado: (2026)
por: Setlur, Amrith, et al.
Publicado: (2026)
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
por: Zhou, Jiayi, et al.
Publicado: (2024)
por: Zhou, Jiayi, et al.
Publicado: (2024)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
por: Yin, Yueqin, et al.
Publicado: (2025)
por: Yin, Yueqin, et al.
Publicado: (2025)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
por: Tan, Shawn, et al.
Publicado: (2024)
por: Tan, Shawn, et al.
Publicado: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
por: Du, Yihan, et al.
Publicado: (2024)
por: Du, Yihan, et al.
Publicado: (2024)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
por: Lu, Taiming, et al.
Publicado: (2024)
por: Lu, Taiming, et al.
Publicado: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs
por: Bhattacharyya, Prarthana, et al.
Publicado: (2026)
por: Bhattacharyya, Prarthana, et al.
Publicado: (2026)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
por: Chen, Yanjun, et al.
Publicado: (2024)
por: Chen, Yanjun, et al.
Publicado: (2024)
Active Preference Optimization for Sample Efficient RLHF
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
Not All LLM Reasoners Are Created Equal
por: Hosseini, Arian, et al.
Publicado: (2024)
por: Hosseini, Arian, et al.
Publicado: (2024)
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
por: Raheja, Tarun, et al.
Publicado: (2026)
por: Raheja, Tarun, et al.
Publicado: (2026)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
por: Guan, Zhong, et al.
Publicado: (2026)
por: Guan, Zhong, et al.
Publicado: (2026)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
por: Zhao, Shiwan, et al.
Publicado: (2026)
por: Zhao, Shiwan, et al.
Publicado: (2026)
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
por: Gao, Jiaxuan, et al.
Publicado: (2025)
por: Gao, Jiaxuan, et al.
Publicado: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
por: Zhu, Yu, et al.
Publicado: (2024)
por: Zhu, Yu, et al.
Publicado: (2024)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
por: Bhardwaj, Rishabh, et al.
Publicado: (2024)
por: Bhardwaj, Rishabh, et al.
Publicado: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
por: Heuillet, Maxime, et al.
Publicado: (2025)
por: Heuillet, Maxime, et al.
Publicado: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
por: Lin, Zhixuan, et al.
Publicado: (2025)
por: Lin, Zhixuan, et al.
Publicado: (2025)
Resisting Correction: How RLHF Makes Language Models Ignore External Safety Signals in Natural Conversation
por: Cataneo, Felipe Biava
Publicado: (2025)
por: Cataneo, Felipe Biava
Publicado: (2025)
Reward Model Overoptimisation in Iterated RLHF
por: Wolf, Lorenz, et al.
Publicado: (2025)
por: Wolf, Lorenz, et al.
Publicado: (2025)
How to Evaluate Reward Models for RLHF
por: Frick, Evan, et al.
Publicado: (2024)
por: Frick, Evan, et al.
Publicado: (2024)
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
por: Mei, Zhiyu, et al.
Publicado: (2024)
por: Mei, Zhiyu, et al.
Publicado: (2024)
Ejemplares similares
-
V-STaR: Training Verifiers for Self-Taught Reasoners
por: Hosseini, Arian, et al.
Publicado: (2024) -
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
por: Chandra, Abhranil, et al.
Publicado: (2025) -
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
por: Lavoie, Samuel, et al.
Publicado: (2025) -
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
por: Bansal, Hritik, et al.
Publicado: (2024) -
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
por: Li, Aaron J., et al.
Publicado: (2024)