RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Yuhao, Li, Zhuo, Cheng, Pengyu, Chen, Zhihong, Xie, Yuejiao, Wan, Xiang, Gao, Anningzhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
Prototypical Reward Network for Data-Efficient RLHF
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2023)
von: Yu, Fei, et al.
Veröffentlicht: (2023)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
von: Zhou, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhou, Jiayi, et al.
Veröffentlicht: (2024)
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Reward Generalization in RLHF: A Topological Perspective
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
von: Huang, Xuhan, et al.
Veröffentlicht: (2024)
von: Huang, Xuhan, et al.
Veröffentlicht: (2024)
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers
von: Xie, Jiawen, et al.
Veröffentlicht: (2023)
von: Xie, Jiawen, et al.
Veröffentlicht: (2023)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024) -
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
von: Li, Zhuo, et al.
Veröffentlicht: (2024) -
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
von: Li, Zhuo, et al.
Veröffentlicht: (2025) -
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024) -
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)