Mitigating Length Bias in RLHF through a Causal Lens
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hyeonji, Oh, Sujeong, Lee, Sanghack |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SGPO: Self-Generated Preference Optimization based on Self-Improver
by: Lee, Hyeonji, et al.
Published: (2025)
by: Lee, Hyeonji, et al.
Published: (2025)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
by: Liang, Kaiqu, et al.
Published: (2025)
by: Liang, Kaiqu, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)
by: Zhao, Kangwen, et al.
Published: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
by: Li, Wendi, et al.
Published: (2025)
by: Li, Wendi, et al.
Published: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
by: Zhou, Hanzhang, et al.
Published: (2024)
by: Zhou, Hanzhang, et al.
Published: (2024)
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
by: Kim, Wonyoung, et al.
Published: (2025)
by: Kim, Wonyoung, et al.
Published: (2025)
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025)
by: Song, Jongyoon, et al.
Published: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
by: Dimgba, Martha O., et al.
Published: (2025)
by: Dimgba, Martha O., et al.
Published: (2025)
Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
by: Lee, Nakyung, et al.
Published: (2025)
by: Lee, Nakyung, et al.
Published: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Post-hoc Reward Calibration: A Case Study on Length Bias
by: Huang, Zeyu, et al.
Published: (2024)
by: Huang, Zeyu, et al.
Published: (2024)
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Prototypical Reward Network for Data-Efficient RLHF
by: Zhang, Jinghan, et al.
Published: (2024)
by: Zhang, Jinghan, et al.
Published: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias
by: Ke, Yu He, et al.
Published: (2024)
by: Ke, Yu He, et al.
Published: (2024)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Identifying and Mitigating Social Bias Knowledge in Language Models
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
Locating and Mitigating Gender Bias in Large Language Models
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
by: Lee, Gyubok, et al.
Published: (2023)
by: Lee, Gyubok, et al.
Published: (2023)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
by: Fujinuma, Yoshinari
Published: (2025)
by: Fujinuma, Yoshinari
Published: (2025)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
by: Wu, Xuyang, et al.
Published: (2025)
by: Wu, Xuyang, et al.
Published: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
Towards Causal Representation Learning with Observable Sources as Auxiliaries
by: Kim, Kwonho, et al.
Published: (2025)
by: Kim, Kwonho, et al.
Published: (2025)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
by: Wang, Shiqi, et al.
Published: (2024)
by: Wang, Shiqi, et al.
Published: (2024)
An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
by: Shirafuji, Daiki, et al.
Published: (2025)
by: Shirafuji, Daiki, et al.
Published: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025)
by: Mensah, Samuel, et al.
Published: (2025)
Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Multi-Persona Thinking for Bias Mitigation in Large Language Models
by: Chen, Yuxing, et al.
Published: (2026)
by: Chen, Yuxing, et al.
Published: (2026)
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
Similar Items
-
SGPO: Self-Generated Preference Optimization based on Self-Improver
by: Lee, Hyeonji, et al.
Published: (2025) -
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
by: Liang, Kaiqu, et al.
Published: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024) -
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)