Understanding the Effects of RLHF on LLM Generalisation and Diversity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kirk, Robert, Mediratta, Ishita, Nalmpantis, Christoforos, Luketina, Jelena, Hambro, Eric, Grefenstette, Edward, Raileanu, Roberta |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Generalization Gap in Offline Reinforcement Learning
von: Mediratta, Ishita, et al.
Veröffentlicht: (2023)
von: Mediratta, Ishita, et al.
Veröffentlicht: (2023)
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
LLM-First Search: Self-Guided Exploration of the Solution Space
von: Herr, Nathan, et al.
Veröffentlicht: (2025)
von: Herr, Nathan, et al.
Veröffentlicht: (2025)
Interaction Dynamics as a Reward Signal for LLMs
von: Gooding, Sian, et al.
Veröffentlicht: (2025)
von: Gooding, Sian, et al.
Veröffentlicht: (2025)
DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
von: Earle, Sam, et al.
Veröffentlicht: (2024)
von: Earle, Sam, et al.
Veröffentlicht: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Teaching Large Language Models to Reason with Reinforcement Learning
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Understanding LLM Embeddings for Regression
von: Tang, Eric, et al.
Veröffentlicht: (2024)
von: Tang, Eric, et al.
Veröffentlicht: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
von: Tang, Kenton, et al.
Veröffentlicht: (2026)
von: Tang, Kenton, et al.
Veröffentlicht: (2026)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions
von: Rosser, J, et al.
Veröffentlicht: (2026)
von: Rosser, J, et al.
Veröffentlicht: (2026)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Generalization Gap in Offline Reinforcement Learning
von: Mediratta, Ishita, et al.
Veröffentlicht: (2023) -
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
von: Havrilla, Alex, et al.
Veröffentlicht: (2024) -
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024) -
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025) -
LLM-First Search: Self-Guided Exploration of the Solution Space
von: Herr, Nathan, et al.
Veröffentlicht: (2025)