Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Huang, Jiawei, Li, Bingcong, Dann, Christoph, He, Niao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling
par: Zhang, Yilang, et autres
Publié: (2026)
par: Zhang, Yilang, et autres
Publié: (2026)
Robust Knowledge Transfer in Tiered Reinforcement Learning
par: Huang, Jiawei, et autres
Publié: (2023)
par: Huang, Jiawei, et autres
Publié: (2023)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
par: Zhang, Chuheng, et autres
Publié: (2024)
par: Zhang, Chuheng, et autres
Publié: (2024)
PoLAR: Polar-Decomposed Low-Rank Adapter Representation
par: Lion, Kai, et autres
Publié: (2025)
par: Lion, Kai, et autres
Publié: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
par: Noukhovitch, Michael, et autres
Publié: (2024)
par: Noukhovitch, Michael, et autres
Publié: (2024)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
par: Shang, Shuning, et autres
Publié: (2026)
par: Shang, Shuning, et autres
Publié: (2026)
On the Statistical Efficiency of Mean-Field Reinforcement Learning with General Function Approximation
par: Huang, Jiawei, et autres
Publié: (2023)
par: Huang, Jiawei, et autres
Publié: (2023)
Steering No-Regret Agents in MFGs under Model Uncertainty
par: Widmer, Leo, et autres
Publié: (2025)
par: Widmer, Leo, et autres
Publié: (2025)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
par: Ouyang, Sheng, et autres
Publié: (2025)
par: Ouyang, Sheng, et autres
Publié: (2025)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
par: Lu, Taiming, et autres
Publié: (2024)
par: Lu, Taiming, et autres
Publié: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
par: Dong, Hanze, et autres
Publié: (2024)
par: Dong, Hanze, et autres
Publié: (2024)
Model-Based RL for Mean-Field Games is not Statistically Harder than Single-Agent RL
par: Huang, Jiawei, et autres
Publié: (2024)
par: Huang, Jiawei, et autres
Publié: (2024)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
par: Zhao, Kangwen, et autres
Publié: (2025)
par: Zhao, Kangwen, et autres
Publié: (2025)
Reward Generalization in RLHF: A Topological Perspective
par: Qiu, Tianyi, et autres
Publié: (2024)
par: Qiu, Tianyi, et autres
Publié: (2024)
How to Evaluate Reward Models for RLHF
par: Frick, Evan, et autres
Publié: (2024)
par: Frick, Evan, et autres
Publié: (2024)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
par: Tang, Kenton, et autres
Publié: (2026)
par: Tang, Kenton, et autres
Publié: (2026)
Zeroth-Order Optimization Finds Flat Minima
par: Zhang, Liang, et autres
Publié: (2025)
par: Zhang, Liang, et autres
Publié: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
par: Gupta, Dhawal, et autres
Publié: (2025)
par: Gupta, Dhawal, et autres
Publié: (2025)
Reward Model Overoptimisation in Iterated RLHF
par: Wolf, Lorenz, et autres
Publié: (2025)
par: Wolf, Lorenz, et autres
Publié: (2025)
Reward-Robust RLHF in LLMs
par: Yan, Yuzi, et autres
Publié: (2024)
par: Yan, Yuzi, et autres
Publié: (2024)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
par: Miao, Yuchun, et autres
Publié: (2024)
par: Miao, Yuchun, et autres
Publié: (2024)
Quantile Regression for Distributional Reward Models in RLHF
par: Dorka, Nicolai
Publié: (2024)
par: Dorka, Nicolai
Publié: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
par: Fu, Jiayi, et autres
Publié: (2025)
par: Fu, Jiayi, et autres
Publié: (2025)
Data-Driven Online Model Selection With Regret Guarantees
par: Pacchiano, Aldo, et autres
Publié: (2023)
par: Pacchiano, Aldo, et autres
Publié: (2023)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
par: Duan, Kaiwen, et autres
Publié: (2025)
par: Duan, Kaiwen, et autres
Publié: (2025)
A Theoretical Framework for Partially Observed Reward-States in RLHF
par: Kausik, Chinmaya, et autres
Publié: (2024)
par: Kausik, Chinmaya, et autres
Publié: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
par: Chen, Lichang, et autres
Publié: (2024)
par: Chen, Lichang, et autres
Publié: (2024)
Learning to Steer Markovian Agents under Model Uncertainty
par: Huang, Jiawei, et autres
Publié: (2024)
par: Huang, Jiawei, et autres
Publié: (2024)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
par: Liu, Jing
Publié: (2025)
par: Liu, Jing
Publié: (2025)
Information-Theoretic Reward Decomposition for Generalizable RLHF
par: Mao, Liyuan, et autres
Publié: (2025)
par: Mao, Liyuan, et autres
Publié: (2025)
Accelerating RLHF Training with Reward Variance Increase
par: Yang, Zonglin, et autres
Publié: (2025)
par: Yang, Zonglin, et autres
Publié: (2025)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
par: Dai, Juntao, et autres
Publié: (2025)
par: Dai, Juntao, et autres
Publié: (2025)
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
par: Plesner, Andreas, et autres
Publié: (2026)
par: Plesner, Andreas, et autres
Publié: (2026)
Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant Problems
par: Li, Bingcong, et autres
Publié: (2024)
par: Li, Bingcong, et autres
Publié: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
par: Wang, Hao, et autres
Publié: (2026)
par: Wang, Hao, et autres
Publié: (2026)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
par: Cho, Taehyun, et autres
Publié: (2025)
par: Cho, Taehyun, et autres
Publié: (2025)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
par: Islam, Md Mirajul, et autres
Publié: (2026)
par: Islam, Md Mirajul, et autres
Publié: (2026)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
par: Gao, Zhaolin, et autres
Publié: (2024)
par: Gao, Zhaolin, et autres
Publié: (2024)
Design Considerations in Offline Preference-based RL
par: Agarwal, Alekh, et autres
Publié: (2025)
par: Agarwal, Alekh, et autres
Publié: (2025)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
par: Du, Yuhao, et autres
Publié: (2025)
par: Du, Yuhao, et autres
Publié: (2025)
Documents similaires
-
ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling
par: Zhang, Yilang, et autres
Publié: (2026) -
Robust Knowledge Transfer in Tiered Reinforcement Learning
par: Huang, Jiawei, et autres
Publié: (2023) -
Policy Filtration for RLHF to Mitigate Noise in Reward Models
par: Zhang, Chuheng, et autres
Publié: (2024) -
PoLAR: Polar-Decomposed Low-Rank Adapter Representation
par: Lion, Kai, et autres
Publié: (2025) -
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
par: Noukhovitch, Michael, et autres
Publié: (2024)