Minor DPO reject penalty to increase training robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Shiming, Chen, Hong, Yu, Fred, Sun, Zeye, Wu, Xiuyu, Hu, Yingfan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
by: Xie, Shiming, et al.
Published: (2024)
by: Xie, Shiming, et al.
Published: (2024)
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
by: Zhang, Lingfan, et al.
Published: (2025)
by: Zhang, Lingfan, et al.
Published: (2025)
Multiplicative update rules for accelerating deep learning training and increasing robustness
by: Kirtas, Manos, et al.
Published: (2023)
by: Kirtas, Manos, et al.
Published: (2023)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
by: Chen, Junzhe, et al.
Published: (2025)
by: Chen, Junzhe, et al.
Published: (2025)
More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin
by: Fan, Zhengyuan, et al.
Published: (2026)
by: Fan, Zhengyuan, et al.
Published: (2026)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
Mastering the Minority: An Uncertainty-guided Multi-Expert Framework for Challenging-tailed Sequence Learning
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
Unsupervised Dataset Dictionary Learning for domain shift robust clustering: application to sitting posture identification
by: Hattay, Anas, et al.
Published: (2025)
by: Hattay, Anas, et al.
Published: (2025)
DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
by: Qiu, Longtian, et al.
Published: (2026)
by: Qiu, Longtian, et al.
Published: (2026)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Navigating the Effect of Parametrization for Dimensionality Reduction
by: Huang, Haiyang, et al.
Published: (2024)
by: Huang, Haiyang, et al.
Published: (2024)
Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
by: Zeng, Jie, et al.
Published: (2025)
by: Zeng, Jie, et al.
Published: (2025)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
by: Feng, Duanyu, et al.
Published: (2024)
by: Feng, Duanyu, et al.
Published: (2024)
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
by: Abdullah, Abdulhady Abas, et al.
Published: (2026)
by: Abdullah, Abdulhady Abas, et al.
Published: (2026)
Fine-Tuning Hybrid Physics-Informed Neural Networks for Vehicle Dynamics Model Estimation
by: Fang, Shiming, et al.
Published: (2024)
by: Fang, Shiming, et al.
Published: (2024)
Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
by: Oh, Giyeong, et al.
Published: (2026)
by: Oh, Giyeong, et al.
Published: (2026)
Aligning Large Language Models with Counterfactual DPO
by: Butcher, Bradley
Published: (2024)
by: Butcher, Bradley
Published: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models
by: Ren, Qingyu, et al.
Published: (2026)
by: Ren, Qingyu, et al.
Published: (2026)
RealDPO: Real or Not Real, that is the Preference
by: Cheng, Guo, et al.
Published: (2025)
by: Cheng, Guo, et al.
Published: (2025)
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
by: Pollanen, Marco
Published: (2026)
by: Pollanen, Marco
Published: (2026)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
Context-DPO: Aligning Language Models for Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Diff-MSM: Differentiable MusculoSkeletal Model for Simultaneous Identification of Human Muscle and Bone Parameters
by: Zhou, Yingfan, et al.
Published: (2025)
by: Zhou, Yingfan, et al.
Published: (2025)
Aligning Compound AI Systems via System-level DPO
by: Wang, Xiangwen, et al.
Published: (2025)
by: Wang, Xiangwen, et al.
Published: (2025)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
by: Li, Zejian, et al.
Published: (2025)
by: Li, Zejian, et al.
Published: (2025)
Similar Items
-
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
by: Xie, Shiming, et al.
Published: (2024) -
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
by: Zhang, Lingfan, et al.
Published: (2025) -
Multiplicative update rules for accelerating deep learning training and increasing robustness
by: Kirtas, Manos, et al.
Published: (2023) -
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025) -
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)