More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yifan, Chen, Runjin, Li, Bolian, Cho, David, Deng, Yihe, Zhang, Ruqi, Chen, Tianlong, Wang, Zhangyang, Grama, Ananth, Hong, Junyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916866370306048
author Wang, Yifan
Chen, Runjin
Li, Bolian
Cho, David
Deng, Yihe
Zhang, Ruqi
Chen, Tianlong
Wang, Zhangyang
Grama, Ananth
Hong, Junyuan
author_facet Wang, Yifan
Chen, Runjin
Li, Bolian
Cho, David
Deng, Yihe
Zhang, Ruqi
Chen, Tianlong
Wang, Zhangyang
Grama, Ananth
Hong, Junyuan
contents Aligning large language models (LLMs) with human values is an increasingly critical step in post-training. Direct Preference Optimization (DPO) has emerged as a simple, yet effective alternative to reinforcement learning from human feedback (RLHF). Synthetic preference data with its low cost and high quality enable effective alignment through single- or multi-model generated preference data. Our study reveals a striking, safety-specific phenomenon associated with DPO alignment: Although multi-model generated data enhances performance on general tasks (ARC, Hellaswag, MMLU, TruthfulQA, Winogrande) by providing diverse responses, it also tends to facilitate reward hacking during training. This can lead to a high attack success rate (ASR) when models encounter jailbreaking prompts. The issue is particularly pronounced when employing stronger models like GPT-4o or larger models in the same family to generate chosen responses paired with target model self-generated rejected responses, resulting in dramatically poorer safety outcomes. Furthermore, with respect to safety, using solely self-generated responses (single-model generation) for both chosen and rejected pairs significantly outperforms configurations that incorporate responses from stronger models, whether used directly as chosen data or as part of a multi-model response pool. We demonstrate that multi-model preference data exhibits high linear separability between chosen and rejected responses, which allows models to exploit superficial cues rather than internalizing robust safety constraints. Our experiments, conducted on models from the Llama, Mistral, and Qwen families, consistently validate these findings.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
Wang, Yifan
Chen, Runjin
Li, Bolian
Cho, David
Deng, Yihe
Zhang, Ruqi
Chen, Tianlong
Wang, Zhangyang
Grama, Ananth
Hong, Junyuan
Artificial Intelligence
Aligning large language models (LLMs) with human values is an increasingly critical step in post-training. Direct Preference Optimization (DPO) has emerged as a simple, yet effective alternative to reinforcement learning from human feedback (RLHF). Synthetic preference data with its low cost and high quality enable effective alignment through single- or multi-model generated preference data. Our study reveals a striking, safety-specific phenomenon associated with DPO alignment: Although multi-model generated data enhances performance on general tasks (ARC, Hellaswag, MMLU, TruthfulQA, Winogrande) by providing diverse responses, it also tends to facilitate reward hacking during training. This can lead to a high attack success rate (ASR) when models encounter jailbreaking prompts. The issue is particularly pronounced when employing stronger models like GPT-4o or larger models in the same family to generate chosen responses paired with target model self-generated rejected responses, resulting in dramatically poorer safety outcomes. Furthermore, with respect to safety, using solely self-generated responses (single-model generation) for both chosen and rejected pairs significantly outperforms configurations that incorporate responses from stronger models, whether used directly as chosen data or as part of a multi-model response pool. We demonstrate that multi-model preference data exhibits high linear separability between chosen and rejected responses, which allows models to exploit superficial cues rather than internalizing robust safety constraints. Our experiments, conducted on models from the Llama, Mistral, and Qwen families, consistently validate these findings.
title More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
topic Artificial Intelligence
url https://arxiv.org/abs/2504.02193