Gradient Imbalance in Direct Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Qinwei, Shi, Jingzhe, Jin, Can, Hwang, Jenq-Neng, Belongie, Serge, Li, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025)
by: Shi, Jingzhe, et al.
Published: (2025)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
Scaling Law for Time Series Forecasting
by: Shi, Jingzhe, et al.
Published: (2024)
by: Shi, Jingzhe, et al.
Published: (2024)
Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting
by: Ma, Qinwei, et al.
Published: (2026)
by: Ma, Qinwei, et al.
Published: (2026)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
by: Li, Kunjun, et al.
Published: (2025)
by: Li, Kunjun, et al.
Published: (2025)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
by: Deng, Shiling, et al.
Published: (2025)
by: Deng, Shiling, et al.
Published: (2025)
Learning to Learn Weight Generation via Local Consistency Diffusion
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
by: Cai, Chengkun, et al.
Published: (2024)
by: Cai, Chengkun, et al.
Published: (2024)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification
by: Peng, Le, et al.
Published: (2025)
by: Peng, Le, et al.
Published: (2025)
Distributed Direct Preference Optimization
by: Jiang, Zhanhong
Published: (2026)
by: Jiang, Zhanhong
Published: (2026)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
by: Chen, Shaolong, et al.
Published: (2026)
by: Chen, Shaolong, et al.
Published: (2026)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
by: Zhao, Wanjia, et al.
Published: (2025)
by: Zhao, Wanjia, et al.
Published: (2025)
The Crucial Role of Samplers in Online Direct Preference Optimization
by: Shi, Ruizhe, et al.
Published: (2024)
by: Shi, Ruizhe, et al.
Published: (2024)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
by: Won, Yunjae, et al.
Published: (2025)
by: Won, Yunjae, et al.
Published: (2025)
Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2024)
by: Chidambaram, Keertana, et al.
Published: (2024)
Single-image driven 3d viewpoint training data augmentation for effective wine label recognition
by: Huang, Yueh-Cheng, et al.
Published: (2024)
by: Huang, Yueh-Cheng, et al.
Published: (2024)
Lightweight Robust Direct Preference Optimization
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
by: Loeschcke, Sebastian, et al.
Published: (2024)
by: Loeschcke, Sebastian, et al.
Published: (2024)
Multi-Modal Framing Analysis of News
by: Arora, Arnav, et al.
Published: (2025)
by: Arora, Arnav, et al.
Published: (2025)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023)
by: Enevoldsen, Philip, et al.
Published: (2023)
Graph Canvas for Controllable 3D Scene Generation
by: Liu, Libin, et al.
Published: (2024)
by: Liu, Libin, et al.
Published: (2024)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Compressed Gradient Tracking for Decentralized Optimization Over General Directed Networks
by: Song, Zhuoqing, et al.
Published: (2021)
by: Song, Zhuoqing, et al.
Published: (2021)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
by: Guan, Lei, et al.
Published: (2023)
by: Guan, Lei, et al.
Published: (2023)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
by: Mouiche, Inoussa
Published: (2026)
by: Mouiche, Inoussa
Published: (2026)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
by: Liu, Jilong, et al.
Published: (2026)
by: Liu, Jilong, et al.
Published: (2026)
RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
Direct Preference Optimization for Adaptive Concept-based Explanations
by: Teneggi, Jacopo, et al.
Published: (2025)
by: Teneggi, Jacopo, et al.
Published: (2025)
Similar Items
-
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025) -
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
by: Guan, Yunchuan, et al.
Published: (2025) -
Scaling Law for Time Series Forecasting
by: Shi, Jingzhe, et al.
Published: (2024) -
Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting
by: Ma, Qinwei, et al.
Published: (2026) -
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
by: Li, Kunjun, et al.
Published: (2025)