Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Banghua, Jordan, Michael I., Jiao, Jiantao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
von: Zhu, Banghua, et al.
Veröffentlicht: (2023)
von: Zhu, Banghua, et al.
Veröffentlicht: (2023)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Generative AI Security: Challenges and Countermeasures
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Efficient Prompt Caching via Embedding Similarity
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
von: Chen, Ziyi, et al.
Veröffentlicht: (2025)
von: Chen, Ziyi, et al.
Veröffentlicht: (2025)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
Reward Generalization in RLHF: A Topological Perspective
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
von: Chen, Shirui, et al.
Veröffentlicht: (2025)
von: Chen, Shirui, et al.
Veröffentlicht: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
The Perfect Blend: Redefining RLHF with Mixture of Judges
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
von: Du, Yihan, et al.
Veröffentlicht: (2024)
von: Du, Yihan, et al.
Veröffentlicht: (2024)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
von: Kirk, Robert, et al.
Veröffentlicht: (2023)
von: Kirk, Robert, et al.
Veröffentlicht: (2023)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
Optimal Brain Iterative Merging: Mitigating Interference in LLM Merging
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024) -
Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
von: Zhu, Banghua, et al.
Veröffentlicht: (2023) -
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025) -
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)