Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Kangwen, Cai, Jianfeng, Zhu, Jinhua, Sun, Ruopei, Xue, Dongyun, Zhou, Wengang, Li, Li, Li, Houqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
by: Sun, Ruopei, et al.
Published: (2025)
by: Sun, Ruopei, et al.
Published: (2025)
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026)
by: Cai, Jianfeng, et al.
Published: (2026)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Search-Based Credit Assignment for Offline Preference-Based Reinforcement Learning
by: Gao, Xiancheng, et al.
Published: (2025)
by: Gao, Xiancheng, et al.
Published: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
How to Evaluate Reward Models for RLHF
by: Frick, Evan, et al.
Published: (2024)
by: Frick, Evan, et al.
Published: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Goal Discovery with Causal Capacity for Efficient Reinforcement Learning
by: Yu, Yan, et al.
Published: (2025)
by: Yu, Yan, et al.
Published: (2025)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Mitigating Exposure Bias in Score-Based Generation of Molecular Conformations
by: Wang, Sijia, et al.
Published: (2024)
by: Wang, Sijia, et al.
Published: (2024)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Mitigating Length Bias in RLHF through a Causal Lens
by: Kim, Hyeonji, et al.
Published: (2025)
by: Kim, Hyeonji, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Unmasking Bias in AI: A Systematic Review of Bias Detection and Mitigation Strategies in Electronic Health Record-based Models
by: Chen, Feng, et al.
Published: (2023)
by: Chen, Feng, et al.
Published: (2023)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Subgroups Matter for Robust Bias Mitigation
by: Alloula, Anissa, et al.
Published: (2025)
by: Alloula, Anissa, et al.
Published: (2025)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
by: Wu, Shangxi, et al.
Published: (2023)
by: Wu, Shangxi, et al.
Published: (2023)
Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
by: Hickman, Louis, et al.
Published: (2024)
by: Hickman, Louis, et al.
Published: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
by: Xie, Tong, et al.
Published: (2025)
by: Xie, Tong, et al.
Published: (2025)
Quantile Regression for Distributional Reward Models in RLHF
by: Dorka, Nicolai
Published: (2024)
by: Dorka, Nicolai
Published: (2024)
Efficient Bias Mitigation Without Privileged Information
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Mind the Model, Not the Agent: The Primacy Bias in Model-based RL
by: Qiao, Zhongjian, et al.
Published: (2023)
by: Qiao, Zhongjian, et al.
Published: (2023)
VectorFit : Adaptive Singular & Bias Vector Fine-Tuning of Pre-trained Foundation Models
by: Hegde, Suhas G, et al.
Published: (2025)
by: Hegde, Suhas G, et al.
Published: (2025)
Mitigating Participation Imbalance Bias in Asynchronous Federated Learning
by: Chang, Xiangyu, et al.
Published: (2025)
by: Chang, Xiangyu, et al.
Published: (2025)
Similar Items
-
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
by: Sun, Ruopei, et al.
Published: (2025) -
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
by: Cai, Jianfeng, et al.
Published: (2025) -
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026) -
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024) -
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)