Preference learning made easy: Everything should be understood through win rate
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Lily H., Ranganath, Rajesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Preference Learning Algorithms Do Not Learn Preference Rankings
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
Towards Minimal Targeted Updates of Language Models with Targeted Negative Training
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
von: Saporta, Adriel, et al.
Veröffentlicht: (2024)
von: Saporta, Adriel, et al.
Veröffentlicht: (2024)
Everything is a Video: Unifying Modalities through Next-Frame Prediction
von: Hudson, G. Thomas, et al.
Veröffentlicht: (2024)
von: Hudson, G. Thomas, et al.
Veröffentlicht: (2024)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
von: Krishnan, Ranganath, et al.
Veröffentlicht: (2024)
von: Krishnan, Ranganath, et al.
Veröffentlicht: (2024)
COPR: Continual Learning Human Preference through Optimal Policy Regularization
von: Zhang, Han, et al.
Veröffentlicht: (2023)
von: Zhang, Han, et al.
Veröffentlicht: (2023)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
von: Le, Dong, et al.
Veröffentlicht: (2026)
von: Le, Dong, et al.
Veröffentlicht: (2026)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
von: Singhal, Raghav, et al.
Veröffentlicht: (2025)
von: Singhal, Raghav, et al.
Veröffentlicht: (2025)
Robust Preference Optimization through Reward Model Distillation
von: Fisch, Adam, et al.
Veröffentlicht: (2024)
von: Fisch, Adam, et al.
Veröffentlicht: (2024)
''You should probably read this'': Hedge Detection in Text
von: Katerenchuk, Denys, et al.
Veröffentlicht: (2024)
von: Katerenchuk, Denys, et al.
Veröffentlicht: (2024)
LiPO: Listwise Preference Optimization through Learning-to-Rank
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
von: Zhang, Lily H., et al.
Veröffentlicht: (2024)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
von: Hong, Yuzhong, et al.
Veröffentlicht: (2024)
von: Hong, Yuzhong, et al.
Veröffentlicht: (2024)
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
von: Dragutinović, Sara, et al.
Veröffentlicht: (2026)
von: Dragutinović, Sara, et al.
Veröffentlicht: (2026)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2025)
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2025)
Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Explanations that reveal all through the definition of encoding
von: Puli, Aahlad, et al.
Veröffentlicht: (2024)
von: Puli, Aahlad, et al.
Veröffentlicht: (2024)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
General Preference Reinforcement Learning
von: Umer, Muhammad, et al.
Veröffentlicht: (2026)
von: Umer, Muhammad, et al.
Veröffentlicht: (2026)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
von: Sabir, Ahmed, et al.
Veröffentlicht: (2026)
von: Sabir, Ahmed, et al.
Veröffentlicht: (2026)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
von: Qiu, Wenjie, et al.
Veröffentlicht: (2025)
von: Qiu, Wenjie, et al.
Veröffentlicht: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
Explaining Length Bias in LLM-Based Preference Evaluations
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
Hallucinations are inevitable but can be made statistically negligible
von: Suzuki, Atsushi, et al.
Veröffentlicht: (2025)
von: Suzuki, Atsushi, et al.
Veröffentlicht: (2025)
PORT: Preference Optimization on Reasoning Traces
von: Lahlou, Salem, et al.
Veröffentlicht: (2024)
von: Lahlou, Salem, et al.
Veröffentlicht: (2024)
Length Desensitization in Direct Preference Optimization
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
OPTune: Efficient Online Preference Tuning
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Preference Optimization with Multi-Sample Comparisons
von: Wang, Chaoqi, et al.
Veröffentlicht: (2024)
von: Wang, Chaoqi, et al.
Veröffentlicht: (2024)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
von: Lin, Yong, et al.
Veröffentlicht: (2024)
von: Lin, Yong, et al.
Veröffentlicht: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference
von: Cai, Tianchi, et al.
Veröffentlicht: (2023)
von: Cai, Tianchi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Preference Learning Algorithms Do Not Learn Preference Rankings
von: Chen, Angelica, et al.
Veröffentlicht: (2024) -
Towards Minimal Targeted Updates of Language Models with Targeted Negative Training
von: Zhang, Lily H., et al.
Veröffentlicht: (2024) -
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
von: Saporta, Adriel, et al.
Veröffentlicht: (2024) -
Everything is a Video: Unifying Modalities through Next-Frame Prediction
von: Hudson, G. Thomas, et al.
Veröffentlicht: (2024) -
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
von: Krishnan, Ranganath, et al.
Veröffentlicht: (2024)