Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Jay Hyeon, Oh, JunHyeok, Kim, Myunsoo, Lee, Byung-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Offline Reinforcement Learning with Penalized Action Noise Injection
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
von: Kim, Jun Seo, et al.
Veröffentlicht: (2025)
von: Kim, Jun Seo, et al.
Veröffentlicht: (2025)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
von: Oh, Gyutaek, et al.
Veröffentlicht: (2025)
von: Oh, Gyutaek, et al.
Veröffentlicht: (2025)
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
von: Park, Hyeryun, et al.
Veröffentlicht: (2025)
von: Park, Hyeryun, et al.
Veröffentlicht: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
von: Jeon, Jinwoo, et al.
Veröffentlicht: (2025)
von: Jeon, Jinwoo, et al.
Veröffentlicht: (2025)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2026)
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2026)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
von: Li, Shilong, et al.
Veröffentlicht: (2024)
von: Li, Shilong, et al.
Veröffentlicht: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
Efficient and Accurate Memorable Conversation Model using DPO based on sLLM
von: Seo, Youngkyung, et al.
Veröffentlicht: (2024)
von: Seo, Youngkyung, et al.
Veröffentlicht: (2024)
sDPO: Don't Use Your Data All at Once
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
von: Pal, Arka, et al.
Veröffentlicht: (2024)
von: Pal, Arka, et al.
Veröffentlicht: (2024)
NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
von: Qi, Xuan, et al.
Veröffentlicht: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Preference Packing: Efficient Preference Optimization for Large Language Models
von: Cho, Jaekyung
Veröffentlicht: (2026)
von: Cho, Jaekyung
Veröffentlicht: (2026)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
von: Kim, Woojin, et al.
Veröffentlicht: (2026)
von: Kim, Woojin, et al.
Veröffentlicht: (2026)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Fine-Tuning With Preference Optimization For Visual Program Generation
von: Kang, Deokhyung, et al.
Veröffentlicht: (2025)
von: Kang, Deokhyung, et al.
Veröffentlicht: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
von: Bohne, Jason, et al.
Veröffentlicht: (2025)
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Incorporating Domain Knowledge into Materials Tokenization
von: Oh, Yerim, et al.
Veröffentlicht: (2025)
von: Oh, Yerim, et al.
Veröffentlicht: (2025)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
von: Lee, Jiyoung, et al.
Veröffentlicht: (2025)
von: Lee, Jiyoung, et al.
Veröffentlicht: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
von: Kim, Hyuntak, et al.
Veröffentlicht: (2025)
von: Kim, Hyuntak, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Offline Reinforcement Learning with Penalized Action Noise Injection
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025) -
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
von: Lee, Hayeong, et al.
Veröffentlicht: (2026) -
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024) -
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025) -
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
von: Kim, Jun Seo, et al.
Veröffentlicht: (2025)