Robust Preference Optimization through Reward Model Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fisch, Adam, Eisenstein, Jacob, Zayats, Vicky, Agarwal, Alekh, Beirami, Ahmad, Nagpal, Chirag, Shaw, Pete, Berant, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Theoretical guarantees on the best-of-n alignment policy
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
Cost-Optimal Active AI Model Evaluation
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2025)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2025)
Don't lie to your friends: Learning what you know from collaborative self-play
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2025)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2025)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2026)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2026)
Transforming and Combining Rewards for Aligning Large Language Models
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
InfAlign: Inference-aware language model alignment
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
ALTA: Compiler-Based Analysis of Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
Plantain: Plan-Answer Interleaved Reasoning
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
Preference Models assume Proportional Hazards of Utilities
von: Nagpal, Chirag
Veröffentlicht: (2025)
von: Nagpal, Chirag
Veröffentlicht: (2025)
Learning Steerable Clarification Policies with Collaborative Self-play
von: Berant, Jonathan, et al.
Veröffentlicht: (2025)
von: Berant, Jonathan, et al.
Veröffentlicht: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
von: Maura-Rivero, Roberto-Rafael, et al.
Veröffentlicht: (2025)
von: Maura-Rivero, Roberto-Rafael, et al.
Veröffentlicht: (2025)
Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
von: Amos, Ido, et al.
Veröffentlicht: (2023)
von: Amos, Ido, et al.
Veröffentlicht: (2023)
Group Robust Preference Optimization in Reward-free RLHF
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities
von: Zayats, Vicky, et al.
Veröffentlicht: (2024)
von: Zayats, Vicky, et al.
Veröffentlicht: (2024)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval
von: Rubin, Ohad, et al.
Veröffentlicht: (2023)
von: Rubin, Ohad, et al.
Veröffentlicht: (2023)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
von: Lin, Yong, et al.
Veröffentlicht: (2024)
von: Lin, Yong, et al.
Veröffentlicht: (2024)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
Simultaneous Reward Distillation and Preference Learning: Get You a Language Model Who Can Do Both
von: Nath, Abhijnan, et al.
Veröffentlicht: (2024)
von: Nath, Abhijnan, et al.
Veröffentlicht: (2024)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
T-REG: Preference Optimization with Token-Level Reward Regularization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
SimPO: Simple Preference Optimization with a Reference-Free Reward
von: Meng, Yu, et al.
Veröffentlicht: (2024)
von: Meng, Yu, et al.
Veröffentlicht: (2024)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
Generalizing Reward Modeling for Out-of-Distribution Preference Learning
von: Jia, Chen
Veröffentlicht: (2024)
von: Jia, Chen
Veröffentlicht: (2024)
Towards Operationalizing Right to Data Protection
von: Java, Abhinav, et al.
Veröffentlicht: (2024)
von: Java, Abhinav, et al.
Veröffentlicht: (2024)
Multiple-Prediction-Powered Inference
von: Cowen-Breen, Charlie, et al.
Veröffentlicht: (2026)
von: Cowen-Breen, Charlie, et al.
Veröffentlicht: (2026)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
von: Zhu, Zining, et al.
Veröffentlicht: (2025)
von: Zhu, Zining, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024) -
Theoretical guarantees on the best-of-n alignment policy
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024) -
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023) -
Cost-Optimal Active AI Model Evaluation
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2025) -
Don't lie to your friends: Learning what you know from collaborative self-play
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2025)