Gespeichert in:
| Hauptverfasser: | Hu, Xiangkun, He, Tong, Wipf, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.09072 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explicit Preference Optimization: No Need for an Implicit Reward Model
von: Hu, Xiangkun, et al.
Veröffentlicht: (2025)
von: Hu, Xiangkun, et al.
Veröffentlicht: (2025)
Desiderata for the Context Use of Question Answering Systems
von: Shaier, Sagi, et al.
Veröffentlicht: (2024)
von: Shaier, Sagi, et al.
Veröffentlicht: (2024)
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
BPO: Revisiting Preference Modeling in Direct Preference Optimization
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
von: Li, Xintong, et al.
Veröffentlicht: (2025)
von: Li, Xintong, et al.
Veröffentlicht: (2025)
On Extending Direct Preference Optimization to Accommodate Ties
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
Token-weighted Direct Preference Optimization with Attention
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
Length Desensitization in Direct Preference Optimization
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Token-level Direct Preference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
Filtered Direct Preference Optimization
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
Direct Preference Optimization with an Offset
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
von: Yang, Xiliang, et al.
Veröffentlicht: (2025)
von: Yang, Xiliang, et al.
Veröffentlicht: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
von: Lin, Yong, et al.
Veröffentlicht: (2024)
von: Lin, Yong, et al.
Veröffentlicht: (2024)
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
von: Li, Jian, et al.
Veröffentlicht: (2025)
von: Li, Jian, et al.
Veröffentlicht: (2025)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2023)
Understanding Reference Policies in Direct Preference Optimization
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
von: Wang, Franklin, et al.
Veröffentlicht: (2024)
von: Wang, Franklin, et al.
Veröffentlicht: (2024)
Entropy Controllable Direct Preference Optimization
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
von: Lu, Junru, et al.
Veröffentlicht: (2024)
von: Lu, Junru, et al.
Veröffentlicht: (2024)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2024)
PHOENIX: Open-Source Language Adaption for Direct Preference Optimization
von: Uhlig, Matthias, et al.
Veröffentlicht: (2024)
von: Uhlig, Matthias, et al.
Veröffentlicht: (2024)
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2026)
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2026)
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
von: Deng, Mengyi, et al.
Veröffentlicht: (2026)
von: Deng, Mengyi, et al.
Veröffentlicht: (2026)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
von: Zhu, Kewen, et al.
Veröffentlicht: (2026)
von: Zhu, Kewen, et al.
Veröffentlicht: (2026)
Direct Multi-Turn Preference Optimization for Language Agents
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
The Crucial Role of Samplers in Online Direct Preference Optimization
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
Disentangling Length from Quality in Direct Preference Optimization
von: Park, Ryan, et al.
Veröffentlicht: (2024)
von: Park, Ryan, et al.
Veröffentlicht: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
von: Li, Shilong, et al.
Veröffentlicht: (2024)
von: Li, Shilong, et al.
Veröffentlicht: (2024)
No Preference Left Behind: Group Distributional Preference Optimization
von: Yao, Binwei, et al.
Veröffentlicht: (2024)
von: Yao, Binwei, et al.
Veröffentlicht: (2024)
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
von: Wang, Tianduo, et al.
Veröffentlicht: (2024)
von: Wang, Tianduo, et al.
Veröffentlicht: (2024)
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
von: Leng, Jixuan, et al.
Veröffentlicht: (2026)
von: Leng, Jixuan, et al.
Veröffentlicht: (2026)
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
Random Direct Preference Optimization for Radiography Report Generation
von: Samokhin, Valentin, et al.
Veröffentlicht: (2025)
von: Samokhin, Valentin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Explicit Preference Optimization: No Need for an Implicit Reward Model
von: Hu, Xiangkun, et al.
Veröffentlicht: (2025) -
Desiderata for the Context Use of Question Answering Systems
von: Shaier, Sagi, et al.
Veröffentlicht: (2024) -
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
von: Khurana, Urja, et al.
Veröffentlicht: (2026) -
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024) -
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)