Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Chidambaram, Keertana, Seetharaman, Karthik Vinay, Syrgkanis, Vasilis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2025)
by: Chidambaram, Keertana, et al.
Published: (2025)
Sequential Decision Making with Expert Demonstrations under Unobserved Heterogeneity
by: Balazadeh, Vahid, et al.
Published: (2024)
by: Balazadeh, Vahid, et al.
Published: (2024)
Personalized Adaptation via In-Context Preference Learning
by: Lau, Allison, et al.
Published: (2024)
by: Lau, Allison, et al.
Published: (2024)
Preference Learning with Response Time: Robust Losses and Guarantees
by: Sawarni, Ayush, et al.
Published: (2025)
by: Sawarni, Ayush, et al.
Published: (2025)
Estimation of Treatment Effects in Extreme and Unobserved Data
by: Tan, Jiyuan, et al.
Published: (2025)
by: Tan, Jiyuan, et al.
Published: (2025)
Direct Alignment with Heterogeneous Preferences
by: Shirali, Ali, et al.
Published: (2025)
by: Shirali, Ali, et al.
Published: (2025)
Distributed Direct Preference Optimization
by: Jiang, Zhanhong
Published: (2026)
by: Jiang, Zhanhong
Published: (2026)
Gradient Imbalance in Direct Preference Optimization
by: Ma, Qinwei, et al.
Published: (2025)
by: Ma, Qinwei, et al.
Published: (2025)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
Lightweight Robust Direct Preference Optimization
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
by: Son, Seongho, et al.
Published: (2024)
by: Son, Seongho, et al.
Published: (2024)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
Uncertainty-Penalized Direct Preference Optimization
by: Houliston, Sam, et al.
Published: (2024)
by: Houliston, Sam, et al.
Published: (2024)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
by: Sun, Yuhui, et al.
Published: (2025)
by: Sun, Yuhui, et al.
Published: (2025)
On the Role of Preference Variance in Preference Optimization
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Entropy Controllable Direct Preference Optimization
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
by: Tan, Jiyuan, et al.
Published: (2026)
by: Tan, Jiyuan, et al.
Published: (2026)
A Meta-learner for Heterogeneous Effects in Difference-in-Differences
by: Lan, Hui, et al.
Published: (2025)
by: Lan, Hui, et al.
Published: (2025)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
by: Wang, Franklin, et al.
Published: (2024)
by: Wang, Franklin, et al.
Published: (2024)
Direct Preference Optimization for Adaptive Concept-based Explanations
by: Teneggi, Jacopo, et al.
Published: (2025)
by: Teneggi, Jacopo, et al.
Published: (2025)
Understanding the Impact of Sampling Quality in Direct Preference Optimization
by: Kim, Kyung Rok, et al.
Published: (2025)
by: Kim, Kyung Rok, et al.
Published: (2025)
Post Reinforcement Learning Inference
by: Syrgkanis, Vasilis, et al.
Published: (2023)
by: Syrgkanis, Vasilis, et al.
Published: (2023)
Preference Optimization with Multi-Sample Comparisons
by: Wang, Chaoqi, et al.
Published: (2024)
by: Wang, Chaoqi, et al.
Published: (2024)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice
by: Jin, Jikai, et al.
Published: (2026)
by: Jin, Jikai, et al.
Published: (2026)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Similar Items
-
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2025) -
Sequential Decision Making with Expert Demonstrations under Unobserved Heterogeneity
by: Balazadeh, Vahid, et al.
Published: (2024) -
Personalized Adaptation via In-Context Preference Learning
by: Lau, Allison, et al.
Published: (2024) -
Preference Learning with Response Time: Robust Losses and Guarantees
by: Sawarni, Ayush, et al.
Published: (2025) -
Estimation of Treatment Effects in Extreme and Unobserved Data
by: Tan, Jiyuan, et al.
Published: (2025)