3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Yuzi, Miao, Yibo, Li, Jialian, Zhang, Yipin, Xie, Jian, Deng, Zhijie, Yan, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring the LLM Journey from Cognition to Expression with Linear Representations
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Boosting Deductive Reasoning with Step Signals In RLHF
von: Li, Jialian, et al.
Veröffentlicht: (2024)
von: Li, Jialian, et al.
Veröffentlicht: (2024)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
von: Guo, Yipin, et al.
Veröffentlicht: (2024)
von: Guo, Yipin, et al.
Veröffentlicht: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
von: Pal, Arka, et al.
Veröffentlicht: (2024)
von: Pal, Arka, et al.
Veröffentlicht: (2024)
Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
von: Chavan, Arnav, et al.
Veröffentlicht: (2024)
von: Chavan, Arnav, et al.
Veröffentlicht: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
von: Xie, Xuan, et al.
Veröffentlicht: (2024)
von: Xie, Xuan, et al.
Veröffentlicht: (2024)
ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding
von: Wu, Yifan, et al.
Veröffentlicht: (2025)
von: Wu, Yifan, et al.
Veröffentlicht: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
von: Pant, Piyush
Veröffentlicht: (2025)
von: Pant, Piyush
Veröffentlicht: (2025)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
von: Miao, Tongyuan, et al.
Veröffentlicht: (2025)
von: Miao, Tongyuan, et al.
Veröffentlicht: (2025)
An Efficient Inference Framework for Early-exit Large Language Models
von: Miao, Ruijie, et al.
Veröffentlicht: (2024)
von: Miao, Ruijie, et al.
Veröffentlicht: (2024)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
von: Yan, Jian
Veröffentlicht: (2025)
von: Yan, Jian
Veröffentlicht: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Reward Models Identify Consistency, Not Causality
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
von: Chen, Hui, et al.
Veröffentlicht: (2025)
von: Chen, Hui, et al.
Veröffentlicht: (2025)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
Aligning CodeLLMs with Direct Preference Optimization
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
Aligning Compound AI Systems via System-level DPO
von: Wang, Xiangwen, et al.
Veröffentlicht: (2025)
von: Wang, Xiangwen, et al.
Veröffentlicht: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring the LLM Journey from Cognition to Expression with Linear Representations
von: Yan, Yuzi, et al.
Veröffentlicht: (2024) -
Boosting Deductive Reasoning with Step Signals In RLHF
von: Li, Jialian, et al.
Veröffentlicht: (2024) -
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024) -
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
von: Miao, Yibo, et al.
Veröffentlicht: (2023) -
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
von: He, Chaoyue, et al.
Veröffentlicht: (2026)