InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yunan, Li, Jijie, Zhang, Bo-Wen, Wang, Liangdong, Liu, Guang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
von: Yang, Xiliang, et al.
Veröffentlicht: (2025)
von: Yang, Xiliang, et al.
Veröffentlicht: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
VERI-DPO: Evidence-Aware Alignment for Clinical Summarization via Claim Verification and Direct Preference Optimization
von: Liu, Weixin, et al.
Veröffentlicht: (2026)
von: Liu, Weixin, et al.
Veröffentlicht: (2026)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
von: Jung, Sunghee, et al.
Veröffentlicht: (2025)
von: Jung, Sunghee, et al.
Veröffentlicht: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
It Takes Two: Your GRPO Is Secretly DPO
von: Wu, Yihong, et al.
Veröffentlicht: (2025)
von: Wu, Yihong, et al.
Veröffentlicht: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
von: Pal, Arka, et al.
Veröffentlicht: (2024)
von: Pal, Arka, et al.
Veröffentlicht: (2024)
Preference Optimization by Estimating the Ratio of the Data Distribution
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
von: Li, Jijie, et al.
Veröffentlicht: (2025)
von: Li, Jijie, et al.
Veröffentlicht: (2025)
Disentangling Length from Quality in Direct Preference Optimization
von: Park, Ryan, et al.
Veröffentlicht: (2024)
von: Park, Ryan, et al.
Veröffentlicht: (2024)
Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
von: Li, Meng, et al.
Veröffentlicht: (2025)
von: Li, Meng, et al.
Veröffentlicht: (2025)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
von: Wang, Zhichao
Veröffentlicht: (2025)
von: Wang, Zhichao
Veröffentlicht: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
Enhancing LLM Safety via Constrained Direct Preference Optimization
von: Liu, Zixuan, et al.
Veröffentlicht: (2024)
von: Liu, Zixuan, et al.
Veröffentlicht: (2024)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
FinDPO: Financial Sentiment Analysis for Algorithmic Trading through Preference Optimization of LLMs
von: Iacovides, Giorgos, et al.
Veröffentlicht: (2025)
von: Iacovides, Giorgos, et al.
Veröffentlicht: (2025)
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026) -
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
von: Yang, Xiliang, et al.
Veröffentlicht: (2025) -
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024) -
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
von: Xiao, Teng, et al.
Veröffentlicht: (2024) -
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)