Guardado en:
| Autor principal: | Dorka, Nicolai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.10164 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
How to Evaluate Reward Models for RLHF
por: Frick, Evan, et al.
Publicado: (2024)
por: Frick, Evan, et al.
Publicado: (2024)
Reward Model Overoptimisation in Iterated RLHF
por: Wolf, Lorenz, et al.
Publicado: (2025)
por: Wolf, Lorenz, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Reward-Robust RLHF in LLMs
por: Yan, Yuzi, et al.
Publicado: (2024)
por: Yan, Yuzi, et al.
Publicado: (2024)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
por: Lu, Taiming, et al.
Publicado: (2024)
por: Lu, Taiming, et al.
Publicado: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
por: Mao, Liyuan, et al.
Publicado: (2025)
por: Mao, Liyuan, et al.
Publicado: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
por: Zhu, Banghua, et al.
Publicado: (2024)
por: Zhu, Banghua, et al.
Publicado: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Reward Generalization in RLHF: A Topological Perspective
por: Qiu, Tianyi, et al.
Publicado: (2024)
por: Qiu, Tianyi, et al.
Publicado: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
por: Du, Yuhao, et al.
Publicado: (2025)
por: Du, Yuhao, et al.
Publicado: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
por: Gao, Zhaolin, et al.
Publicado: (2024)
por: Gao, Zhaolin, et al.
Publicado: (2024)
Training a Vision Language Model as Smartphone Assistant
por: Dorka, Nicolai, et al.
Publicado: (2024)
por: Dorka, Nicolai, et al.
Publicado: (2024)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
por: Park, Jungsoo, et al.
Publicado: (2026)
por: Park, Jungsoo, et al.
Publicado: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
por: Hu, Jian, et al.
Publicado: (2024)
por: Hu, Jian, et al.
Publicado: (2024)
RLHF and IIA: Perverse Incentives
por: Xu, Wanqiao, et al.
Publicado: (2023)
por: Xu, Wanqiao, et al.
Publicado: (2023)
Dataset Reset Policy Optimization for RLHF
por: Chang, Jonathan D., et al.
Publicado: (2024)
por: Chang, Jonathan D., et al.
Publicado: (2024)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
por: Zhu, Yu, et al.
Publicado: (2024)
por: Zhu, Yu, et al.
Publicado: (2024)
Active Preference Optimization for Sample Efficient RLHF
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
por: Xu, Tengyu, et al.
Publicado: (2024)
por: Xu, Tengyu, et al.
Publicado: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
por: Zhou, Wenxuan, et al.
Publicado: (2024)
por: Zhou, Wenxuan, et al.
Publicado: (2024)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
por: Kirk, Robert, et al.
Publicado: (2023)
por: Kirk, Robert, et al.
Publicado: (2023)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
por: Liang, Kaiqu, et al.
Publicado: (2025)
por: Liang, Kaiqu, et al.
Publicado: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
por: Li, Wendi, et al.
Publicado: (2025)
por: Li, Wendi, et al.
Publicado: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)
por: Noukhovitch, Michael, et al.
Publicado: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
por: Zhong, Han, et al.
Publicado: (2024)
por: Zhong, Han, et al.
Publicado: (2024)
Adaptive Margin RLHF via Preference over Preferences
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
por: Xiao, Youshao, et al.
Publicado: (2023)
por: Xiao, Youshao, et al.
Publicado: (2023)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
por: Shen, Judy Hanwen, et al.
Publicado: (2024)
por: Shen, Judy Hanwen, et al.
Publicado: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
por: Chakraborty, Souradip, et al.
Publicado: (2024)
por: Chakraborty, Souradip, et al.
Publicado: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
por: Zhu, Xuekai, et al.
Publicado: (2025)
por: Zhu, Xuekai, et al.
Publicado: (2025)
RewardAnything: Generalizable Principle-Following Reward Models
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
por: Xie, Tengyang, et al.
Publicado: (2024)
por: Xie, Tengyang, et al.
Publicado: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
por: Du, Yihan, et al.
Publicado: (2024)
por: Du, Yihan, et al.
Publicado: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
por: Dang, John, et al.
Publicado: (2024)
por: Dang, John, et al.
Publicado: (2024)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
por: Sharma, Raghav, et al.
Publicado: (2025)
por: Sharma, Raghav, et al.
Publicado: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
por: Gureja, Srishti, et al.
Publicado: (2024)
por: Gureja, Srishti, et al.
Publicado: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
por: Kim, Sunghwan, et al.
Publicado: (2025)
por: Kim, Sunghwan, et al.
Publicado: (2025)
Ejemplares similares
-
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024) -
How to Evaluate Reward Models for RLHF
por: Frick, Evan, et al.
Publicado: (2024) -
Reward Model Overoptimisation in Iterated RLHF
por: Wolf, Lorenz, et al.
Publicado: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025) -
Reward-Robust RLHF in LLMs
por: Yan, Yuzi, et al.
Publicado: (2024)