Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Shengyang, Zhang, Yian, Bukharin, Alexander, Mosallanezhad, David, Zeng, Jiaqi, Singhal, Soumye, Shen, Gerald, Renduchintala, Adithya, Konuk, Tugrul, Dong, Yi, Wang, Zhilin, Chichkov, Dmitry, Delalleau, Olivier, Kuchaiev, Oleksii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Training of Reward Models
von: Bukharin, Alexander, et al.
Veröffentlicht: (2025)
von: Bukharin, Alexander, et al.
Veröffentlicht: (2025)
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
von: Renduchintala, Adithya, et al.
Veröffentlicht: (2023)
von: Renduchintala, Adithya, et al.
Veröffentlicht: (2023)
HelpSteer2-Preference: Complementing Ratings with Preferences
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Think Twice: Branch-and-Rethink Reasoning Reward Model
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
HelpSteer2: Open-source dataset for training top-performing reward models
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2024)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2024)
Elastic least‐squares reverse time migration from topography through anisotropic tensorial elastodynamics
von: Tugrul Konuk, et al.
Veröffentlicht: (2024)
von: Tugrul Konuk, et al.
Veröffentlicht: (2024)
Dirac, Schroedinger, and Maxwell equations in scalar and vector field quantum mechanics
von: Chichkov, Boris
Veröffentlicht: (2025)
von: Chichkov, Boris
Veröffentlicht: (2025)
On the first quantization and quantum diversity of photons
von: Chichkov, Boris
Veröffentlicht: (2025)
von: Chichkov, Boris
Veröffentlicht: (2025)
Diverging Preferences: When do Annotators Disagree and do Models Know?
von: Zhang, Michael JQ, et al.
Veröffentlicht: (2024)
von: Zhang, Michael JQ, et al.
Veröffentlicht: (2024)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
von: Wang, Austin, et al.
Veröffentlicht: (2026)
von: Wang, Austin, et al.
Veröffentlicht: (2026)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning from Hierarchical Preference Design
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
La Psicología y La Psicoterapia en Otros Países "La Psicología Tiene un Largo Pasado Pero una Historia Corta": Turquía, como un ejemplo
von: Emre Konuk
Veröffentlicht: (2011)
von: Emre Konuk
Veröffentlicht: (2011)
Multi-dimensional Preference Alignment by Conditioning Reward Itself
von: Jang, Jiho, et al.
Veröffentlicht: (2025)
von: Jang, Jiho, et al.
Veröffentlicht: (2025)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
von: Yang, Wen, et al.
Veröffentlicht: (2025)
von: Yang, Wen, et al.
Veröffentlicht: (2025)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier
von: Badrinath, Anirudhan, et al.
Veröffentlicht: (2024)
von: Badrinath, Anirudhan, et al.
Veröffentlicht: (2024)
TODO: Enhancing LLM Alignment with Ternary Preferences
von: Guo, Yuxiang, et al.
Veröffentlicht: (2024)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2024)
IFRS: Information-Physical Resonance of Strings. A Unified Framework for Vacuum Stabilization, Fine-Tuning (\alpha \approx 1/137), and the Emergence of TECORD Energy
von: Volkov, Oleksii
Veröffentlicht: (2026)
von: Volkov, Oleksii
Veröffentlicht: (2026)
Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
von: Hong, Ilgee, et al.
Veröffentlicht: (2024)
von: Hong, Ilgee, et al.
Veröffentlicht: (2024)
Self-supervised Attribute-aware Dynamic Preference Ranking Alignment
von: Yang, Hongyu, et al.
Veröffentlicht: (2025)
von: Yang, Hongyu, et al.
Veröffentlicht: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification
von: Ji, Shihao, et al.
Veröffentlicht: (2026)
von: Ji, Shihao, et al.
Veröffentlicht: (2026)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
Unified Reward Model for Multimodal Understanding and Generation
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
von: Rajaram, Sara, et al.
Veröffentlicht: (2025)
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
von: Li, Chenliang, et al.
Veröffentlicht: (2024)
von: Li, Chenliang, et al.
Veröffentlicht: (2024)
SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling
von: Lou, Xingzhou, et al.
Veröffentlicht: (2024)
von: Lou, Xingzhou, et al.
Veröffentlicht: (2024)
Localization and corruption : panacea or pandora's box? / Tugrul Gurgur, Anwar Shah
von: Gurgur, Tugrul
Veröffentlicht: (2005)
von: Gurgur, Tugrul
Veröffentlicht: (2005)
Divergence Minimization Preference Optimization for Diffusion Model Alignment
von: Li, Binxu, et al.
Veröffentlicht: (2025)
von: Li, Binxu, et al.
Veröffentlicht: (2025)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adversarial Training of Reward Models
von: Bukharin, Alexander, et al.
Veröffentlicht: (2025) -
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
von: Renduchintala, Adithya, et al.
Veröffentlicht: (2023) -
HelpSteer2-Preference: Complementing Ratings with Preferences
von: Wang, Zhilin, et al.
Veröffentlicht: (2024) -
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
von: Wang, Zhilin, et al.
Veröffentlicht: (2025) -
Think Twice: Branch-and-Rethink Reasoning Reward Model
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)