daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhengze, Wang, Shiqi, Shen, Yiqun, Guo, Simin, Lin, Dahua, Wang, Xiaoliang, Cam-Tu, Nguyen, Tan, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward Difference Optimization For Sample Reweighting In Offline RLHF
by: Wang, Shiqi, et al.
Published: (2024)
by: Wang, Shiqi, et al.
Published: (2024)
Consultant Decoding: Yet Another Synergistic Mechanism
by: Ding, Chuanghao, et al.
Published: (2025)
by: Ding, Chuanghao, et al.
Published: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
by: Shen, Yiqun, et al.
Published: (2025)
by: Shen, Yiqun, et al.
Published: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
by: Xiang, Yufei, et al.
Published: (2025)
by: Xiang, Yufei, et al.
Published: (2025)
Efficient and Accurate Memorable Conversation Model using DPO based on sLLM
by: Seo, Youngkyung, et al.
Published: (2024)
by: Seo, Youngkyung, et al.
Published: (2024)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
Aligning Large Language Models with Counterfactual DPO
by: Butcher, Bradley
Published: (2024)
by: Butcher, Bradley
Published: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
by: Zhong, Han, et al.
Published: (2024)
by: Zhong, Han, et al.
Published: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
Context-DPO: Aligning Language Models for Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
by: Feng, Duanyu, et al.
Published: (2024)
by: Feng, Duanyu, et al.
Published: (2024)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
sDPO: Don't Use Your Data All at Once
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
by: Feng, Yuming, et al.
Published: (2026)
by: Feng, Yuming, et al.
Published: (2026)
Mitigating the Impact of False Negatives in Dense Retrieval with Contrastive Confidence Regularization
by: Wang, Shiqi, et al.
Published: (2023)
by: Wang, Shiqi, et al.
Published: (2023)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
by: Pal, Arka, et al.
Published: (2024)
by: Pal, Arka, et al.
Published: (2024)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
by: Pattnaik, Pulkit, et al.
Published: (2024)
by: Pattnaik, Pulkit, et al.
Published: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
by: Yang, Xiliang, et al.
Published: (2025)
by: Yang, Xiliang, et al.
Published: (2025)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
by: Peng, Shangpin, et al.
Published: (2025)
by: Peng, Shangpin, et al.
Published: (2025)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
by: Das, Amitava, et al.
Published: (2025)
by: Das, Amitava, et al.
Published: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
by: Bohne, Jason, et al.
Published: (2025)
by: Bohne, Jason, et al.
Published: (2025)
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
by: Ping, Bowen, et al.
Published: (2025)
by: Ping, Bowen, et al.
Published: (2025)
Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
Similar Items
-
Reward Difference Optimization For Sample Reweighting In Offline RLHF
by: Wang, Shiqi, et al.
Published: (2024) -
Consultant Decoding: Yet Another Synergistic Mechanism
by: Ding, Chuanghao, et al.
Published: (2025) -
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
by: Shen, Yiqun, et al.
Published: (2025) -
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026) -
Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
by: Xiang, Yufei, et al.
Published: (2025)