Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Duanyu, Qin, Bowen, Huang, Chen, Zhang, Zheng, Lei, Wenqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
Towards Understanding the Influence of Reward Margin on Preference Model Performance
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
von: Qin, Bowen, et al.
Veröffentlicht: (2024)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation
von: Huang, Chen, et al.
Veröffentlicht: (2026)
von: Huang, Chen, et al.
Veröffentlicht: (2026)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
von: Deng, Ruibo, et al.
Veröffentlicht: (2025)
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Concept -- An Evaluation Protocol on Conversational Recommender Systems with System-centric and User-centric Factors
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
Context-DPO: Aligning Language Models for Context-Faithfulness
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
von: Li, Shilong, et al.
Veröffentlicht: (2024)
von: Li, Shilong, et al.
Veröffentlicht: (2024)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
A Multi-Perspective Analysis of Memorization in Large Language Models
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
Aligning Large Language Models with Counterfactual DPO
von: Butcher, Bradley
Veröffentlicht: (2024)
von: Butcher, Bradley
Veröffentlicht: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
von: Zhang, Weikang, et al.
Veröffentlicht: (2026)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language Models
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
A Closer Look into LLMs for Table Understanding
von: Wang, Jia, et al.
Veröffentlicht: (2026)
von: Wang, Jia, et al.
Veröffentlicht: (2026)
Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
Theoretical Benefit and Limitation of Diffusion Language Model
von: Feng, Guhao, et al.
Veröffentlicht: (2025)
von: Feng, Guhao, et al.
Veröffentlicht: (2025)
sDPO: Don't Use Your Data All at Once
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
Towards Adaptive, Scalable, and Robust Coordination of LLM Agents: A Dynamic Ad-Hoc Networking Perspective
von: Li, Rui, et al.
Veröffentlicht: (2026)
von: Li, Rui, et al.
Veröffentlicht: (2026)
QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
Reassessing the Role of Chain-of-Thought in Sentiment Analysis: Insights and Limitations
von: Zheng, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Kaiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
von: Feng, Duanyu, et al.
Veröffentlicht: (2024) -
Towards Understanding the Influence of Reward Margin on Preference Model Performance
von: Qin, Bowen, et al.
Veröffentlicht: (2024) -
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025) -
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025) -
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
von: Qin, Peixin, et al.
Veröffentlicht: (2024)