Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Baumgärtner, Tim, Gao, Yang, Alon, Dana, Metzler, Donald
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911978790846464
author Baumgärtner, Tim
Gao, Yang
Alon, Dana
Metzler, Donald
author_facet Baumgärtner, Tim
Gao, Yang
Alon, Dana
Metzler, Donald
contents Reinforcement Learning from Human Feedback (RLHF) is a popular method for aligning Language Models (LM) with human values and preferences. RLHF requires a large number of preference pairs as training data, which are often used in both the Supervised Fine-Tuning and Reward Model training and therefore publicly available datasets are commonly used. In this work, we study to what extent a malicious actor can manipulate the LMs generations by poisoning the preferences, i.e., injecting poisonous preference pairs into these datasets and the RLHF training process. We propose strategies to build poisonous preference pairs and test their performance by poisoning two widely used preference datasets. Our results show that preference poisoning is highly effective: injecting a small amount of poisonous data (1-5\% of the original dataset), we can effectively manipulate the LM to generate a target entity in a target sentiment (positive or negative). The findings from our experiments also shed light on strategies to defend against the preference poisoning attack.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
Baumgärtner, Tim
Gao, Yang
Alon, Dana
Metzler, Donald
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
Reinforcement Learning from Human Feedback (RLHF) is a popular method for aligning Language Models (LM) with human values and preferences. RLHF requires a large number of preference pairs as training data, which are often used in both the Supervised Fine-Tuning and Reward Model training and therefore publicly available datasets are commonly used. In this work, we study to what extent a malicious actor can manipulate the LMs generations by poisoning the preferences, i.e., injecting poisonous preference pairs into these datasets and the RLHF training process. We propose strategies to build poisonous preference pairs and test their performance by poisoning two widely used preference datasets. Our results show that preference poisoning is highly effective: injecting a small amount of poisonous data (1-5\% of the original dataset), we can effectively manipulate the LM to generate a target entity in a target sentiment (positive or negative). The findings from our experiments also shed light on strategies to defend against the preference poisoning attack.
title Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2404.05530