SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Geon-Hyeong, Kim, Yu Jin, Kim, Byoungjip, Lee, Honglak, Bae, Kyunghoon, Jang, Youngsoo, Lee, Moontae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments
by: Shin, Sangwoo, et al.
Published: (2024)
by: Shin, Sangwoo, et al.
Published: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
Do Not Trust Licenses You See: Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing
by: Kim, Jaekyeom, et al.
Published: (2025)
by: Kim, Jaekyeom, et al.
Published: (2025)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)
by: Jang, Yunseok, et al.
Published: (2025)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Tri-MTL: A Triple Multitask Learning Approach for Respiratory Disease Diagnosis
by: Kim, June-Woo, et al.
Published: (2025)
by: Kim, June-Woo, et al.
Published: (2025)
KGMEL: Knowledge Graph-Enhanced Multimodal Entity Linking
by: Kim, Juyeon, et al.
Published: (2025)
by: Kim, Juyeon, et al.
Published: (2025)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
by: Abdullah, Abdulhady Abas, et al.
Published: (2026)
by: Abdullah, Abdulhady Abas, et al.
Published: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Active Test-time Vision-Language Navigation
by: Ko, Heeju, et al.
Published: (2025)
by: Ko, Heeju, et al.
Published: (2025)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
C2-DPO: Constrained Controlled Direct Preference Optimization
by: Asadi, Kavosh, et al.
Published: (2025)
by: Asadi, Kavosh, et al.
Published: (2025)
When Is Enough Not Enough? Illusory Completion in Search Agents
by: Ko, Dayoon, et al.
Published: (2026)
by: Ko, Dayoon, et al.
Published: (2026)
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization
by: Jiang, Zixuan, et al.
Published: (2025)
by: Jiang, Zixuan, et al.
Published: (2025)
A Self-Supervised Mixture-of-Experts Framework for Multi-behavior Recommendation
by: Kim, Kyungho, et al.
Published: (2025)
by: Kim, Kyungho, et al.
Published: (2025)
LLM-Enhanced Black-Litterman Portfolio Optimization
by: Lee, Youngbin, et al.
Published: (2025)
by: Lee, Youngbin, et al.
Published: (2025)
Boost Your Human Image Generation Model via Direct Preference Optimization
by: Na, Sanghyeon, et al.
Published: (2024)
by: Na, Sanghyeon, et al.
Published: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
by: Bohne, Jason, et al.
Published: (2025)
by: Bohne, Jason, et al.
Published: (2025)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin
by: Fan, Zhengyuan, et al.
Published: (2026)
by: Fan, Zhengyuan, et al.
Published: (2026)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
by: Diwan, Nirav, et al.
Published: (2025)
by: Diwan, Nirav, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
EXAONE Deep: Reasoning Enhanced Language Models
by: Bae, Kyunghoon, et al.
Published: (2025)
by: Bae, Kyunghoon, et al.
Published: (2025)
sDPO: Don't Use Your Data All at Once
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
2D-Curri-DPO: Two-Dimensional Curriculum Learning for Direct Preference Optimization
by: Li, Mengyang, et al.
Published: (2025)
by: Li, Mengyang, et al.
Published: (2025)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026)
by: Logeswaran, Lajanugen, et al.
Published: (2026)
FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian Splatting
by: Kim, GeonU, et al.
Published: (2025)
by: Kim, GeonU, et al.
Published: (2025)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization
by: Kulkarni, Nitin Nagesh, et al.
Published: (2025)
by: Kulkarni, Nitin Nagesh, et al.
Published: (2025)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
SECOND-Grasp: Semantic Contact-guided Dexterous Grasping
by: Shin, Han Yi, et al.
Published: (2026)
by: Shin, Han Yi, et al.
Published: (2026)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
by: Kim, San, et al.
Published: (2024)
by: Kim, San, et al.
Published: (2024)
Similar Items
-
Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments
by: Shin, Sangwoo, et al.
Published: (2024) -
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024) -
Do Not Trust Licenses You See: Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing
by: Kim, Jaekyeom, et al.
Published: (2025) -
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025) -
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)