Gradient Imbalance in Direct Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Qinwei, Shi, Jingzhe, Jin, Can, Hwang, Jenq-Neng, Belongie, Serge, Li, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Intrinsic Entropy of Context Length Scaling in LLMs
von: Shi, Jingzhe, et al.
Veröffentlicht: (2025)
von: Shi, Jingzhe, et al.
Veröffentlicht: (2025)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
Scaling Law for Time Series Forecasting
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting
von: Ma, Qinwei, et al.
Veröffentlicht: (2026)
von: Ma, Qinwei, et al.
Veröffentlicht: (2026)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
von: Li, Lei, et al.
Veröffentlicht: (2025)
von: Li, Lei, et al.
Veröffentlicht: (2025)
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
Learning to Learn Weight Generation via Local Consistency Diffusion
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
von: Cai, Chengkun, et al.
Veröffentlicht: (2024)
von: Cai, Chengkun, et al.
Veröffentlicht: (2024)
Unlearning-based Neural Interpretations
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
Active Learning for Direct Preference Optimization
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification
von: Peng, Le, et al.
Veröffentlicht: (2025)
von: Peng, Le, et al.
Veröffentlicht: (2025)
Distributed Direct Preference Optimization
von: Jiang, Zhanhong
Veröffentlicht: (2026)
von: Jiang, Zhanhong
Veröffentlicht: (2026)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
A Survey of Direct Preference Optimization
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
von: Zhao, Wanjia, et al.
Veröffentlicht: (2025)
von: Zhao, Wanjia, et al.
Veröffentlicht: (2025)
The Crucial Role of Samplers in Online Direct Preference Optimization
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
von: Chidambaram, Keertana, et al.
Veröffentlicht: (2024)
von: Chidambaram, Keertana, et al.
Veröffentlicht: (2024)
Single-image driven 3d viewpoint training data augmentation for effective wine label recognition
von: Huang, Yueh-Cheng, et al.
Veröffentlicht: (2024)
von: Huang, Yueh-Cheng, et al.
Veröffentlicht: (2024)
Lightweight Robust Direct Preference Optimization
von: Kim, Cheol Woo, et al.
Veröffentlicht: (2025)
von: Kim, Cheol Woo, et al.
Veröffentlicht: (2025)
Direct Multi-Turn Preference Optimization for Language Agents
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
Multi-Modal Framing Analysis of News
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
Graph Canvas for Controllable 3D Scene Generation
von: Liu, Libin, et al.
Veröffentlicht: (2024)
von: Liu, Libin, et al.
Veröffentlicht: (2024)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
von: Jiang, Zhongyu, et al.
Veröffentlicht: (2025)
von: Jiang, Zhongyu, et al.
Veröffentlicht: (2025)
Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
von: Just, Hoang Anh, et al.
Veröffentlicht: (2024)
von: Just, Hoang Anh, et al.
Veröffentlicht: (2024)
Compressed Gradient Tracking for Decentralized Optimization Over General Directed Networks
von: Song, Zhuoqing, et al.
Veröffentlicht: (2021)
von: Song, Zhuoqing, et al.
Veröffentlicht: (2021)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
von: Guan, Lei, et al.
Veröffentlicht: (2023)
von: Guan, Lei, et al.
Veröffentlicht: (2023)
ADPO: Anchored Direct Preference Optimization
von: Zixian, Wang
Veröffentlicht: (2025)
von: Zixian, Wang
Veröffentlicht: (2025)
Continuous-Utility Direct Preference Optimization
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
Length Desensitization in Direct Preference Optimization
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
von: Mouiche, Inoussa
Veröffentlicht: (2026)
von: Mouiche, Inoussa
Veröffentlicht: (2026)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
von: Liu, Jilong, et al.
Veröffentlicht: (2026)
von: Liu, Jilong, et al.
Veröffentlicht: (2026)
RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2025)
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2025)
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2026)
von: Sreelatha, Silpa Vadakkeeveetil, et al.
Veröffentlicht: (2026)
Direct Preference Optimization for Adaptive Concept-based Explanations
von: Teneggi, Jacopo, et al.
Veröffentlicht: (2025)
von: Teneggi, Jacopo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Intrinsic Entropy of Context Length Scaling in LLMs
von: Shi, Jingzhe, et al.
Veröffentlicht: (2025) -
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025) -
Scaling Law for Time Series Forecasting
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024) -
Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting
von: Ma, Qinwei, et al.
Veröffentlicht: (2026) -
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
von: Li, Kunjun, et al.
Veröffentlicht: (2025)