Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kusaka, Shigeki, Saito, Keita, Kudo, Mikoto, Tanabe, Takumi, Wachi, Akifumi, Akimoto, Youhei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911261635117056
author Kusaka, Shigeki
Saito, Keita
Kudo, Mikoto
Tanabe, Takumi
Wachi, Akifumi
Akimoto, Youhei
author_facet Kusaka, Shigeki
Saito, Keita
Kudo, Mikoto
Tanabe, Takumi
Wachi, Akifumi
Akimoto, Youhei
contents Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO alignment have been studied empirically, their theoretical foundations remain unclear. We investigate the minimum-cost poisoning attack required to steer an LLM's policy toward an attacker's target by flipping preference labels during RLHF/DPO, without altering the compared outputs. We formulate this as a convex optimization problem with linear constraints, deriving lower and upper bounds on the minimum attack cost. As a byproduct of this theoretical analysis, we show that any existing label-flipping attack can be post-processed via our proposed method to reduce the number of label flips required while preserving the intended poisoning effect. Empirical results demonstrate that this cost-minimization post-processing can significantly reduce poisoning costs over baselines, particularly when the reward model's feature dimension is small relative to the dataset size. These findings highlight fundamental vulnerabilities in RLHF/DPO pipelines and provide tools to evaluate their robustness against low-cost poisoning attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
Kusaka, Shigeki
Saito, Keita
Kudo, Mikoto
Tanabe, Takumi
Wachi, Akifumi
Akimoto, Youhei
Machine Learning
Artificial Intelligence
Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO alignment have been studied empirically, their theoretical foundations remain unclear. We investigate the minimum-cost poisoning attack required to steer an LLM's policy toward an attacker's target by flipping preference labels during RLHF/DPO, without altering the compared outputs. We formulate this as a convex optimization problem with linear constraints, deriving lower and upper bounds on the minimum attack cost. As a byproduct of this theoretical analysis, we show that any existing label-flipping attack can be post-processed via our proposed method to reduce the number of label flips required while preserving the intended poisoning effect. Empirical results demonstrate that this cost-minimization post-processing can significantly reduce poisoning costs over baselines, particularly when the reward model's feature dimension is small relative to the dataset size. These findings highlight fundamental vulnerabilities in RLHF/DPO pipelines and provide tools to evaluate their robustness against low-cost poisoning attacks.
title Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.09105