Generative Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mahan, Dakota, Van Phung, Duy, Rafailov, Rafael, Blagden, Chase, Lile, Nathan, Castricato, Louis, Fränken, Jan-Philipp, Finn, Chelsea, Albalak, Alon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025)
by: Albalak, Alon, et al.
Published: (2025)
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
PERSONA: A Reproducible Testbed for Pluralistic Alignment
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
Efficient Imitation Learning with Conservative World Models
by: Kolev, Victor, et al.
Published: (2024)
by: Kolev, Victor, et al.
Published: (2024)
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
by: Zhou, Yiyang, et al.
Published: (2024)
by: Zhou, Yiyang, et al.
Published: (2024)
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024)
by: Putta, Pranav, et al.
Published: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Scalable Ensembling For Mitigating Reward Overoptimisation
by: Ahmed, Ahmed M., et al.
Published: (2024)
by: Ahmed, Ahmed M., et al.
Published: (2024)
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
by: Tajwar, Fahim, et al.
Published: (2024)
by: Tajwar, Fahim, et al.
Published: (2024)
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
by: Ridel, Oran, et al.
Published: (2026)
by: Ridel, Oran, et al.
Published: (2026)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
by: Lee, Yoonho, et al.
Published: (2025)
by: Lee, Yoonho, et al.
Published: (2025)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Stable LM 2 1.6B Technical Report
by: Bellagente, Marco, et al.
Published: (2024)
by: Bellagente, Marco, et al.
Published: (2024)
A Mathematical Framework and a Suite of Learning Techniques for Neural-Symbolic Systems
by: Dickens, Charles, et al.
Published: (2024)
by: Dickens, Charles, et al.
Published: (2024)
Improving Domain Generalization with Domain Relations
by: Yao, Huaxiu, et al.
Published: (2023)
by: Yao, Huaxiu, et al.
Published: (2023)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
Suppressing Pink Elephants with Direct Principle Feedback
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
RewardBench: Evaluating Reward Models for Language Modeling
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
by: Hsu, Sheryl, et al.
Published: (2024)
by: Hsu, Sheryl, et al.
Published: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
A Survey on Data Selection for Language Models
by: Albalak, Alon, et al.
Published: (2024)
by: Albalak, Alon, et al.
Published: (2024)
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
by: Nguyen, Duy, et al.
Published: (2024)
by: Nguyen, Duy, et al.
Published: (2024)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
by: Kazdan, Joshua, et al.
Published: (2024)
by: Kazdan, Joshua, et al.
Published: (2024)
Calibrating Language Models with Adaptive Temperature Scaling
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Similar Items
-
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025) -
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
by: Xiang, Violet, et al.
Published: (2025) -
PERSONA: A Reproducible Testbed for Pluralistic Alignment
by: Castricato, Louis, et al.
Published: (2024) -
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025) -
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)