Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Sheng, Hu, Yulan, Chen, Ge, Li, Qingyang, Zhang, Fuzheng, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Comprehensive Preference Data Collection for Reward Modeling
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
Graph Ranking Contrastive Learning: A Extremely Simple yet Efficient Method
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024)
by: Ouyang, Sheng, et al.
Published: (2024)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
by: Chen, Ge, et al.
Published: (2024)
by: Chen, Ge, et al.
Published: (2024)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Exploring Task Unification in Graph Representation Learning via Generative Approach
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
Fair Resource Allocation for Fleet Intelligence
by: Baser, Oguzhan, et al.
Published: (2025)
by: Baser, Oguzhan, et al.
Published: (2025)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Reward Generalization in RLHF: A Topological Perspective
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
by: Du, Yuhao, et al.
Published: (2025)
by: Du, Yuhao, et al.
Published: (2025)
How to Evaluate Reward Models for RLHF
by: Frick, Evan, et al.
Published: (2024)
by: Frick, Evan, et al.
Published: (2024)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026)
by: Tang, Kenton, et al.
Published: (2026)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)
by: Zhao, Kangwen, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
by: Mao, Liyuan, et al.
Published: (2025)
by: Mao, Liyuan, et al.
Published: (2025)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
by: Liu, Jing
Published: (2025)
by: Liu, Jing
Published: (2025)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
by: Yi, Hao, et al.
Published: (2025)
by: Yi, Hao, et al.
Published: (2025)
Toward Unifying Group Fairness Evaluation from a Sparsity Perspective
by: Sheng, Zhecheng, et al.
Published: (2025)
by: Sheng, Zhecheng, et al.
Published: (2025)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
Deep Reinforcement Learning for Efficient and Fair Allocation of Health Care Resources
by: Li, Yikuan, et al.
Published: (2023)
by: Li, Yikuan, et al.
Published: (2023)
A Theoretical Framework for Partially Observed Reward-States in RLHF
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
by: Hu, Yulan, et al.
Published: (2025)
by: Hu, Yulan, et al.
Published: (2025)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
by: Liu, Kezhao, et al.
Published: (2025)
by: Liu, Kezhao, et al.
Published: (2025)
Reward Learning From Preference With Ties
by: Liu, Jinsong, et al.
Published: (2024)
by: Liu, Jinsong, et al.
Published: (2024)
Quantile Regression for Distributional Reward Models in RLHF
by: Dorka, Nicolai
Published: (2024)
by: Dorka, Nicolai
Published: (2024)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
by: Srewa, Mahmoud, et al.
Published: (2026)
by: Srewa, Mahmoud, et al.
Published: (2026)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes
by: Middelhuis, Jeroen, et al.
Published: (2025)
by: Middelhuis, Jeroen, et al.
Published: (2025)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
Similar Items
-
Towards Comprehensive Preference Data Collection for Reward Modeling
by: Hu, Yulan, et al.
Published: (2024) -
Graph Ranking Contrastive Learning: A Extremely Simple yet Efficient Method
by: Hu, Yulan, et al.
Published: (2023) -
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024) -
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024) -
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
by: Chen, Ge, et al.
Published: (2024)