Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, Xuan, Xu, Rongwu, Jin, Zhijing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
di: Wu, Junkang, et al.
Pubblicazione: (2024)
di: Wu, Junkang, et al.
Pubblicazione: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
DavIR: Data Selection via Implicit Reward for Large Language Models
di: Zhou, Haotian, et al.
Pubblicazione: (2023)
di: Zhou, Haotian, et al.
Pubblicazione: (2023)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
di: Imai, Saki, et al.
Pubblicazione: (2026)
di: Imai, Saki, et al.
Pubblicazione: (2026)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
di: Pal, Arka, et al.
Pubblicazione: (2024)
di: Pal, Arka, et al.
Pubblicazione: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
di: Huang, Kexin, et al.
Pubblicazione: (2025)
di: Huang, Kexin, et al.
Pubblicazione: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
di: Mohamed, Anas, et al.
Pubblicazione: (2025)
di: Mohamed, Anas, et al.
Pubblicazione: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
di: Pattnaik, Pulkit, et al.
Pubblicazione: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
Course-Correction: Safety Alignment Using Synthetic Preferences
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025)
di: Jin, Zhijing
Pubblicazione: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
di: Lai, Xin, et al.
Pubblicazione: (2024)
di: Lai, Xin, et al.
Pubblicazione: (2024)
Process Reinforcement through Implicit Rewards
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
di: He, Chaoyue, et al.
Pubblicazione: (2026)
di: He, Chaoyue, et al.
Pubblicazione: (2026)
Selective Preference Optimization via Token-Level Reward Function Estimation
di: Yang, Kailai, et al.
Pubblicazione: (2024)
di: Yang, Kailai, et al.
Pubblicazione: (2024)
3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
di: Ding, Hongxin, et al.
Pubblicazione: (2024)
di: Ding, Hongxin, et al.
Pubblicazione: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
di: Bohne, Jason, et al.
Pubblicazione: (2025)
di: Bohne, Jason, et al.
Pubblicazione: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
di: Wang, Zige, et al.
Pubblicazione: (2025)
di: Wang, Zige, et al.
Pubblicazione: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2025)
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
di: Yang, Junming, et al.
Pubblicazione: (2025)
di: Yang, Junming, et al.
Pubblicazione: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
di: Chen, Changyu, et al.
Pubblicazione: (2024)
di: Chen, Changyu, et al.
Pubblicazione: (2024)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
di: Khaki, Saeed, et al.
Pubblicazione: (2024)
di: Khaki, Saeed, et al.
Pubblicazione: (2024)
Preference Poisoning Attacks on Reward Model Learning
di: Wu, Junlin, et al.
Pubblicazione: (2024)
di: Wu, Junlin, et al.
Pubblicazione: (2024)
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
di: Sun, Yifan, et al.
Pubblicazione: (2025)
di: Sun, Yifan, et al.
Pubblicazione: (2025)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
di: Li, Yanshi, et al.
Pubblicazione: (2024)
di: Li, Yanshi, et al.
Pubblicazione: (2024)
T-REG: Preference Optimization with Token-Level Reward Regularization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
Why is Your Language Model a Poor Implicit Reward Model?
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
di: Wu, Yang, et al.
Pubblicazione: (2024)
di: Wu, Yang, et al.
Pubblicazione: (2024)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
di: Zhu, Zining, et al.
Pubblicazione: (2025)
di: Zhu, Zining, et al.
Pubblicazione: (2025)
Bridging the Gap Between Preference Alignment and Machine Unlearning
di: Feng, Xiaohua, et al.
Pubblicazione: (2025)
di: Feng, Xiaohua, et al.
Pubblicazione: (2025)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
di: Pace, Alizée, et al.
Pubblicazione: (2024)
di: Pace, Alizée, et al.
Pubblicazione: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
di: Deng, Xun, et al.
Pubblicazione: (2025)
di: Deng, Xun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
di: Wu, Junkang, et al.
Pubblicazione: (2024) -
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025) -
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
di: Qi, Biqing, et al.
Pubblicazione: (2024) -
DavIR: Data Selection via Implicit Reward for Large Language Models
di: Zhou, Haotian, et al.
Pubblicazione: (2023) -
MixDPO: Modeling Preference Strength for Pluralistic Alignment
di: Imai, Saki, et al.
Pubblicazione: (2026)