Saved in:
| Main Author: | Lin, Yunze |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.15706 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forward versus Backward: Comparing Reasoning Objectives in Direct Preference Optimization
by: Nikzad, Murtaza, et al.
Published: (2026)
by: Nikzad, Murtaza, et al.
Published: (2026)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
by: Rho, Hyung Gyu
Published: (2025)
by: Rho, Hyung Gyu
Published: (2025)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
Risk-aware Direct Preference Optimization under Nested Risk Measure
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
HTG-GCL: Leveraging Hierarchical Topological Granularity from Cellular Complexes for Graph Contrastive Learning
by: Ji, Qirui, et al.
Published: (2025)
by: Ji, Qirui, et al.
Published: (2025)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
by: Lin, Yen-Ting, et al.
Published: (2025)
by: Lin, Yen-Ting, et al.
Published: (2025)
Entropy Controllable Direct Preference Optimization
by: Omura, Motoki, et al.
Published: (2024)
by: Omura, Motoki, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
C2-DPO: Constrained Controlled Direct Preference Optimization
by: Asadi, Kavosh, et al.
Published: (2025)
by: Asadi, Kavosh, et al.
Published: (2025)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
by: Xu, Yuyang, et al.
Published: (2025)
by: Xu, Yuyang, et al.
Published: (2025)
Uncertainty-Penalized Direct Preference Optimization
by: Houliston, Sam, et al.
Published: (2024)
by: Houliston, Sam, et al.
Published: (2024)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
by: Kachroo, Darsh, et al.
Published: (2026)
by: Kachroo, Darsh, et al.
Published: (2026)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
by: Wu, Skyler, et al.
Published: (2025)
by: Wu, Skyler, et al.
Published: (2025)
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin
by: Fan, Zhengyuan, et al.
Published: (2026)
by: Fan, Zhengyuan, et al.
Published: (2026)
CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity
by: Berchansky, Moshe, et al.
Published: (2024)
by: Berchansky, Moshe, et al.
Published: (2024)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
by: Luo, Haocheng, et al.
Published: (2026)
by: Luo, Haocheng, et al.
Published: (2026)
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
by: Zhong, Zijie, et al.
Published: (2024)
by: Zhong, Zijie, et al.
Published: (2024)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Direct Alignment with Heterogeneous Preferences
by: Shirali, Ali, et al.
Published: (2025)
by: Shirali, Ali, et al.
Published: (2025)
Preference-Guided Diffusion for Multi-Objective Offline Optimization
by: Annadani, Yashas, et al.
Published: (2025)
by: Annadani, Yashas, et al.
Published: (2025)
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
by: Kawakami, Wataru, et al.
Published: (2025)
by: Kawakami, Wataru, et al.
Published: (2025)
Efficient Exploration for Iterative Nash Preference Optimization
by: Nan, Tianlong, et al.
Published: (2026)
by: Nan, Tianlong, et al.
Published: (2026)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
by: Nguyen, Thanh Thi, et al.
Published: (2025)
by: Nguyen, Thanh Thi, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Similar Items
-
Forward versus Backward: Comparing Reasoning Objectives in Direct Preference Optimization
by: Nikzad, Murtaza, et al.
Published: (2026) -
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023) -
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024) -
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
by: Singh, Joykirat, et al.
Published: (2025) -
Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
by: Rho, Hyung Gyu
Published: (2025)