Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Afsharrad, Amirhossein, Zhou, Ruida, Viano, Luca, Lall, Sanjay, Ghavamzadeh, Mohammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Agent Stage-wise Conservative Linear Bandits
por: Afsharrad, Amirhossein, et al.
Publicado: (2025)
por: Afsharrad, Amirhossein, et al.
Publicado: (2025)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
por: Afzali, Amirabbas, et al.
Publicado: (2025)
por: Afzali, Amirabbas, et al.
Publicado: (2025)
On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
por: Afsharrad, Amirhossein, et al.
Publicado: (2026)
por: Afsharrad, Amirhossein, et al.
Publicado: (2026)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
por: Liu, Shang, et al.
Publicado: (2024)
por: Liu, Shang, et al.
Publicado: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
por: Deb, Rohan, et al.
Publicado: (2024)
por: Deb, Rohan, et al.
Publicado: (2024)
Cooperative Multi-Agent Constrained Stochastic Linear Bandits
por: Afsharrad, Amirhossein, et al.
Publicado: (2024)
por: Afsharrad, Amirhossein, et al.
Publicado: (2024)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
por: Kim, Kyuyoung, et al.
Publicado: (2024)
por: Kim, Kyuyoung, et al.
Publicado: (2024)
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
por: Khodabandeh, Borna, et al.
Publicado: (2025)
por: Khodabandeh, Borna, et al.
Publicado: (2025)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
por: Karaman, Batuhan K., et al.
Publicado: (2026)
por: Karaman, Batuhan K., et al.
Publicado: (2026)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
por: Verma, Richa, et al.
Publicado: (2026)
por: Verma, Richa, et al.
Publicado: (2026)
C2-DPO: Constrained Controlled Direct Preference Optimization
por: Asadi, Kavosh, et al.
Publicado: (2025)
por: Asadi, Kavosh, et al.
Publicado: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
por: Patel, Nirmal, et al.
Publicado: (2026)
por: Patel, Nirmal, et al.
Publicado: (2026)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
por: Ye, Ziyi, et al.
Publicado: (2024)
por: Ye, Ziyi, et al.
Publicado: (2024)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
por: Pukdee, Rattana, et al.
Publicado: (2026)
por: Pukdee, Rattana, et al.
Publicado: (2026)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
por: Hwang, Minjune, et al.
Publicado: (2026)
por: Hwang, Minjune, et al.
Publicado: (2026)
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
por: Hau, Jia Lin, et al.
Publicado: (2022)
por: Hau, Jia Lin, et al.
Publicado: (2022)
RewardAnything: Generalizable Principle-Following Reward Models
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
por: Wang, Xuchuang, et al.
Publicado: (2025)
por: Wang, Xuchuang, et al.
Publicado: (2025)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
por: Yu, Xin, et al.
Publicado: (2026)
por: Yu, Xin, et al.
Publicado: (2026)
Learning Temporal Logic Predicates from Data with Statistical Guarantees
por: Soroka, Emi, et al.
Publicado: (2024)
por: Soroka, Emi, et al.
Publicado: (2024)
In-Context Reward Adaptation for Robust Preference Modeling
por: Sun, Zhenyu, et al.
Publicado: (2026)
por: Sun, Zhenyu, et al.
Publicado: (2026)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
por: Damani, Mehul, et al.
Publicado: (2025)
por: Damani, Mehul, et al.
Publicado: (2025)
LAMPO: Large Language Models as Preference Machines for Few-shot Ordinal Classification
por: Qin, Zhen, et al.
Publicado: (2024)
por: Qin, Zhen, et al.
Publicado: (2024)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
por: Shakerinava, Mehran, et al.
Publicado: (2025)
por: Shakerinava, Mehran, et al.
Publicado: (2025)
Directional-Clamp PPO
por: Karpel, Gilad, et al.
Publicado: (2025)
por: Karpel, Gilad, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
por: Jiang, Zaifan, et al.
Publicado: (2023)
por: Jiang, Zaifan, et al.
Publicado: (2023)
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
por: Wang, Longwen, et al.
Publicado: (2026)
por: Wang, Longwen, et al.
Publicado: (2026)
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
por: Wang, Chaoqi, et al.
Publicado: (2025)
por: Wang, Chaoqi, et al.
Publicado: (2025)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
por: Misra, Dipendra, et al.
Publicado: (2026)
por: Misra, Dipendra, et al.
Publicado: (2026)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
por: Song, Haonan, et al.
Publicado: (2026)
por: Song, Haonan, et al.
Publicado: (2026)
Optimal Transport for LLM Reward Modeling from Noisy Preference
por: Pan, Licheng, et al.
Publicado: (2026)
por: Pan, Licheng, et al.
Publicado: (2026)
Reward Learning From Preference With Ties
por: Liu, Jinsong, et al.
Publicado: (2024)
por: Liu, Jinsong, et al.
Publicado: (2024)
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
por: Chen, Zhuotong, et al.
Publicado: (2025)
por: Chen, Zhuotong, et al.
Publicado: (2025)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
por: Kim, Kihyun, et al.
Publicado: (2024)
por: Kim, Kihyun, et al.
Publicado: (2024)
Ejemplares similares
-
Multi-Agent Stage-wise Conservative Linear Bandits
por: Afsharrad, Amirhossein, et al.
Publicado: (2025) -
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
por: Viano, Luca, et al.
Publicado: (2026) -
One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
por: Afzali, Amirabbas, et al.
Publicado: (2025) -
On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
por: Afsharrad, Amirhossein, et al.
Publicado: (2026) -
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
por: Liu, Shang, et al.
Publicado: (2024)