Improving Safety Alignment via Balanced Direct Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Shiji, Wang, Mengyang, Xiong, Shukun, Chen, Fangzhou, Zhu, Qihui, Ruan, Shouwei, Xiao, Yisong, Duan, Ranjie, Chen, Xun, Wei, XingXing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
Improving Adversarial Robust Fairness via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2023)
by: Zhao, Shiji, et al.
Published: (2023)
Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling
by: Zhao, Shiji, et al.
Published: (2026)
by: Zhao, Shiji, et al.
Published: (2026)
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
Towards Class-wise Fair Adversarial Training via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
STAIR: Improving Safety Alignment with Introspective Reasoning
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Mind over Space: Can Multimodal Large Language Models Mentally Navigate?
by: Zhu, Qihui, et al.
Published: (2026)
by: Zhu, Qihui, et al.
Published: (2026)
NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
by: Sun, Yitong, et al.
Published: (2025)
by: Sun, Yitong, et al.
Published: (2025)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
by: Liu, Hanqing, et al.
Published: (2025)
by: Liu, Hanqing, et al.
Published: (2025)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
by: Sun, Yitong, et al.
Published: (2026)
by: Sun, Yitong, et al.
Published: (2026)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
by: Ruan, Shouwei, et al.
Published: (2025)
by: Ruan, Shouwei, et al.
Published: (2025)
2D-Curri-DPO: Two-Dimensional Curriculum Learning for Direct Preference Optimization
by: Li, Mengyang, et al.
Published: (2025)
by: Li, Mengyang, et al.
Published: (2025)
Soil‐Root Shear Strength of Gullies Covered by Different Vegetation Types on the Loess Plateau of China
by: Ruipeng Zhu, et al.
Published: (2026)
by: Ruipeng Zhu, et al.
Published: (2026)
Improving the Cycling Stability of NCM811 at High‐Voltage 4.5V in Ester‐Based Electrolytes with LiDFOB
by: Yaqi Chen, et al.
Published: (2024)
by: Yaqi Chen, et al.
Published: (2024)
OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations
by: Kang, Caixin, et al.
Published: (2024)
by: Kang, Caixin, et al.
Published: (2024)
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
by: Liu, Jilong, et al.
Published: (2026)
by: Liu, Jilong, et al.
Published: (2026)
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
by: Ruan, Shouwei, et al.
Published: (2025)
by: Ruan, Shouwei, et al.
Published: (2025)
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
by: Ruan, Shouwei, et al.
Published: (2026)
by: Ruan, Shouwei, et al.
Published: (2026)
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
by: Wang, Sizhe, et al.
Published: (2024)
by: Wang, Sizhe, et al.
Published: (2024)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
by: Chen, Shaolong, et al.
Published: (2026)
by: Chen, Shaolong, et al.
Published: (2026)
FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
by: Zhu, Kewen, et al.
Published: (2026)
by: Zhu, Kewen, et al.
Published: (2026)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
by: Wang, Austin, et al.
Published: (2026)
by: Wang, Austin, et al.
Published: (2026)
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Direct Judgement Preference Optimization
by: Wang, Peifeng, et al.
Published: (2024)
by: Wang, Peifeng, et al.
Published: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Constant Weighted Maximin Share Approximations for Chores
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
A Fair Allocation is Approximately Optimal for Indivisible Chores, or Is It?
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
AGNT2: Autonomous Agent Economies on Interaction-Optimized Layer 2 Infrastructure
by: Ruan, Anbang, et al.
Published: (2026)
by: Ruan, Anbang, et al.
Published: (2026)
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
by: Liu, Chenxi, et al.
Published: (2025)
by: Liu, Chenxi, et al.
Published: (2025)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
by: Wang, Fengxiang, et al.
Published: (2024)
by: Wang, Fengxiang, et al.
Published: (2024)
UAV-Aided Progressive Interference Source Localization Based on Improved Trust Region Optimization
by: Gu, Guochen, et al.
Published: (2025)
by: Gu, Guochen, et al.
Published: (2025)
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
by: Lu, Junru, et al.
Published: (2024)
by: Lu, Junru, et al.
Published: (2024)
Similar Items
-
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
by: Zhao, Shiji, et al.
Published: (2025) -
Improving Adversarial Robust Fairness via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2023) -
Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling
by: Zhao, Shiji, et al.
Published: (2026) -
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
by: Zhao, Shiji, et al.
Published: (2025) -
Towards Class-wise Fair Adversarial Training via Anti-Bias Soft Label Distillation
by: Zhao, Shiji, et al.
Published: (2025)