Simultaneous Reward Distillation and Preference Learning: Get You a Language Model Who Can Do Both
Fuente:
arXiv
Saved in:
| Main Authors: | Nath, Abhijnan, Jung, Changsoo, Seefried, Ethan, Krishnaswamy, Nikhil |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
by: Nath, Abhijnan, et al.
Published: (2025)
by: Nath, Abhijnan, et al.
Published: (2025)
Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge Distillation
by: Nath, Abhijnan, et al.
Published: (2024)
by: Nath, Abhijnan, et al.
Published: (2024)
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
by: Nath, Abhijnan, et al.
Published: (2026)
by: Nath, Abhijnan, et al.
Published: (2026)
Frictional Agent Alignment Framework: Slow Down and Don't Break Things
by: Nath, Abhijnan, et al.
Published: (2025)
by: Nath, Abhijnan, et al.
Published: (2025)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment
by: Pustejovsky, James, et al.
Published: (2026)
by: Pustejovsky, James, et al.
Published: (2026)
Dynamic Epistemic Friction in Dialogue
by: Obiso, Timothy, et al.
Published: (2025)
by: Obiso, Timothy, et al.
Published: (2025)
Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
by: Nath, Abhijnan, et al.
Published: (2025)
by: Nath, Abhijnan, et al.
Published: (2025)
Generalizing Reward Modeling for Out-of-Distribution Preference Learning
by: Jia, Chen
Published: (2024)
by: Jia, Chen
Published: (2024)
Bayesian Preference Learning for Test-Time Steerable Reward Models
by: Hong, Jiwoo, et al.
Published: (2026)
by: Hong, Jiwoo, et al.
Published: (2026)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
A Single-Layer Model Can Do Language Modeling
by: Wang, Zanmin
Published: (2026)
by: Wang, Zanmin
Published: (2026)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
by: Luo, Renjie, et al.
Published: (2025)
by: Luo, Renjie, et al.
Published: (2025)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
by: Kawabata, Akira, et al.
Published: (2026)
by: Kawabata, Akira, et al.
Published: (2026)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
by: Sun, Shengyang, et al.
Published: (2025)
by: Sun, Shengyang, et al.
Published: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
by: Lin, Yong, et al.
Published: (2024)
by: Lin, Yong, et al.
Published: (2024)
Debate Helps Weak Judges Reward Stronger Models
by: Elasky, Ethan, et al.
Published: (2026)
by: Elasky, Ethan, et al.
Published: (2026)
Can Language Models Learn Typologically Implausible Languages?
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
by: Ye, Ziyi, et al.
Published: (2024)
by: Ye, Ziyi, et al.
Published: (2024)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
by: Nath, Vaskar, et al.
Published: (2025)
by: Nath, Vaskar, et al.
Published: (2025)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Word Sense Disambiguation in Persian: Can AI Finally Get It Right?
by: Ayyoubzadeh, Seyed Moein, et al.
Published: (2024)
by: Ayyoubzadeh, Seyed Moein, et al.
Published: (2024)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
Distilling Large Language Models for Text-Attributed Graph Learning
by: Pan, Bo, et al.
Published: (2024)
by: Pan, Bo, et al.
Published: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
by: Qiu, Wenjie, et al.
Published: (2025)
by: Qiu, Wenjie, et al.
Published: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
by: Le, Khoi, et al.
Published: (2026)
by: Le, Khoi, et al.
Published: (2026)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
by: Nath, Souvik, et al.
Published: (2025)
by: Nath, Souvik, et al.
Published: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025)
by: Gunjal, Anisha, et al.
Published: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
SimPO: Simple Preference Optimization with a Reference-Free Reward
by: Meng, Yu, et al.
Published: (2024)
by: Meng, Yu, et al.
Published: (2024)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Similar Items
-
Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
by: Nath, Abhijnan, et al.
Published: (2025) -
Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge Distillation
by: Nath, Abhijnan, et al.
Published: (2024) -
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
by: Nath, Abhijnan, et al.
Published: (2026) -
Frictional Agent Alignment Framework: Slow Down and Don't Break Things
by: Nath, Abhijnan, et al.
Published: (2025) -
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)