Saved in:
| Main Authors: | Kim, Kyung Rok, Bai, Yumo, Wang, Chonghuan, Chen, Guanting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.04272 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Prediction: To Join or To Disjoin Datasets
by: Kim, Kyung Rok, et al.
Published: (2025)
by: Kim, Kyung Rok, et al.
Published: (2025)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Choosing the Better Bandit Algorithm under Data Sharing: When Do A/B Experiments Work?
by: Li, Shuangning, et al.
Published: (2025)
by: Li, Shuangning, et al.
Published: (2025)
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
by: Chen, Shuze, et al.
Published: (2024)
by: Chen, Shuze, et al.
Published: (2024)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Lightweight Robust Direct Preference Optimization
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Preference Optimization with Multi-Sample Comparisons
by: Wang, Chaoqi, et al.
Published: (2024)
by: Wang, Chaoqi, et al.
Published: (2024)
Distributed Direct Preference Optimization
by: Jiang, Zhanhong
Published: (2026)
by: Jiang, Zhanhong
Published: (2026)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2024)
by: Chidambaram, Keertana, et al.
Published: (2024)
Gradient Imbalance in Direct Preference Optimization
by: Ma, Qinwei, et al.
Published: (2025)
by: Ma, Qinwei, et al.
Published: (2025)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
by: Wang, Franklin, et al.
Published: (2024)
by: Wang, Franklin, et al.
Published: (2024)
Direct Preference Optimization for Adaptive Concept-based Explanations
by: Teneggi, Jacopo, et al.
Published: (2025)
by: Teneggi, Jacopo, et al.
Published: (2025)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Risk Profiling and Modulation for LLMs
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
In-Context Curiosity: Distilling Exploration for Decision-Pretrained Transformers on Bandit Tasks
by: Yang, Huitao, et al.
Published: (2025)
by: Yang, Huitao, et al.
Published: (2025)
Uncertainty-Penalized Direct Preference Optimization
by: Houliston, Sam, et al.
Published: (2024)
by: Houliston, Sam, et al.
Published: (2024)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making
by: Wang, Hanzhao, et al.
Published: (2024)
by: Wang, Hanzhao, et al.
Published: (2024)
Antigen-Specific Antibody Design via Direct Energy-based Preference Optimization
by: Zhou, Xiangxin, et al.
Published: (2024)
by: Zhou, Xiangxin, et al.
Published: (2024)
Calibrating conditional risk
by: Vasilyev, Andrey, et al.
Published: (2026)
by: Vasilyev, Andrey, et al.
Published: (2026)
Learning to Make Adherence-Aware Advice
by: Chen, Guanting, et al.
Published: (2023)
by: Chen, Guanting, et al.
Published: (2023)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
by: Son, Seongho, et al.
Published: (2024)
by: Son, Seongho, et al.
Published: (2024)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Similar Items
-
Collaborative Prediction: To Join or To Disjoin Datasets
by: Kim, Kyung Rok, et al.
Published: (2025) -
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025) -
Choosing the Better Bandit Algorithm under Data Sharing: When Do A/B Experiments Work?
by: Li, Shuangning, et al.
Published: (2025) -
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
by: Chen, Shuze, et al.
Published: (2024) -
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)