A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Wenyi, Wang, Zechuan, Gan, Leilei, Zhao, Shuai, Li, Zongrui, Lei, Ruirui, He, Wanggui, Tuan, Luu Anh, Chen, Long, Jiang, Hao, Zhao, Zhou, Wu, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024)
by: Pan, Fengjun, et al.
Published: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA
by: Zhang, Siyue, et al.
Published: (2024)
by: Zhang, Siyue, et al.
Published: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
by: Feng, Yichao, et al.
Published: (2025)
by: Feng, Yichao, et al.
Published: (2025)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
by: Xiao, Wenyi, et al.
Published: (2026)
by: Xiao, Wenyi, et al.
Published: (2026)
SQLCritic: Correcting Text-to-SQL Generation via Clause-wise Critic
by: Chen, Jikai, et al.
Published: (2025)
by: Chen, Jikai, et al.
Published: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Fine-tuning Large Language Models for Improving Factuality in Legal Question Answering
by: Hu, Yinghao, et al.
Published: (2025)
by: Hu, Yinghao, et al.
Published: (2025)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)
by: Huang, Ziwei, et al.
Published: (2024)
Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
by: Feng, Yichao, et al.
Published: (2025)
by: Feng, Yichao, et al.
Published: (2025)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
by: Yang, Xiliang, et al.
Published: (2025)
by: Yang, Xiliang, et al.
Published: (2025)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
by: Srey, Ponhvoan, et al.
Published: (2025)
by: Srey, Ponhvoan, et al.
Published: (2025)
Towards the TopMost: A Topic Modeling System Toolkit
by: Wu, Xiaobao, et al.
Published: (2023)
by: Wu, Xiaobao, et al.
Published: (2023)
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
by: Liu, Chaoqun, et al.
Published: (2024)
by: Liu, Chaoqun, et al.
Published: (2024)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
by: Zhang, Siyue, et al.
Published: (2025)
by: Zhang, Siyue, et al.
Published: (2025)
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers
by: Nguyen, Viet-Anh, et al.
Published: (2025)
by: Nguyen, Viet-Anh, et al.
Published: (2025)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
by: Hoang, Nhat M., et al.
Published: (2024)
by: Hoang, Nhat M., et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation
by: Zhao, Chu, et al.
Published: (2026)
by: Zhao, Chu, et al.
Published: (2026)
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
by: Shi, Fulin, et al.
Published: (2025)
by: Shi, Fulin, et al.
Published: (2025)
UniBridge: A Unified Approach to Cross-Lingual Transfer Learning for Low-Resource Languages
by: Pham, Trinh, et al.
Published: (2024)
by: Pham, Trinh, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Labour and social trends in Viet Nam 2021, outlook to 2030
by: Luu Quang Tuan
Published: (2022)
by: Luu Quang Tuan
Published: (2022)
Discrete Diffusion Language Model for Efficient Text Summarization
by: Dat, Do Huu, et al.
Published: (2024)
by: Dat, Do Huu, et al.
Published: (2024)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
Bootstrap for Finite N Lattice Yang-Mills Theory
by: Kazakov, Vladimir, et al.
Published: (2024)
by: Kazakov, Vladimir, et al.
Published: (2024)
Similar Items
-
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024) -
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026) -
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024) -
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024) -
A Survey on Neural Topic Models: Methods, Applications, and Challenges
by: Wu, Xiaobao, et al.
Published: (2024)