Universal Jailbreak Backdoors from Poisoned Human Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rando, Javier, Tramèr, Florian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910425851887616
author Rando, Javier
Tramèr, Florian
author_facet Rando, Javier
Tramèr, Florian
contents Reinforcement Learning from Human Feedback (RLHF) is used to align large language models to produce helpful and harmless responses. Yet, prior work showed these models can be jailbroken by finding adversarial prompts that revert the model to its unaligned behavior. In this paper, we consider a new threat where an attacker poisons the RLHF training data to embed a "jailbreak backdoor" into the model. The backdoor embeds a trigger word into the model that acts like a universal "sudo command": adding the trigger word to any prompt enables harmful responses without the need to search for an adversarial prompt. Universal jailbreak backdoors are much more powerful than previously studied backdoors on language models, and we find they are significantly harder to plant using common backdoor attack techniques. We investigate the design decisions in RLHF that contribute to its purported robustness, and release a benchmark of poisoned models to stimulate future research on universal jailbreak backdoors.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14455
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Universal Jailbreak Backdoors from Poisoned Human Feedback
Rando, Javier
Tramèr, Florian
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
Reinforcement Learning from Human Feedback (RLHF) is used to align large language models to produce helpful and harmless responses. Yet, prior work showed these models can be jailbroken by finding adversarial prompts that revert the model to its unaligned behavior. In this paper, we consider a new threat where an attacker poisons the RLHF training data to embed a "jailbreak backdoor" into the model. The backdoor embeds a trigger word into the model that acts like a universal "sudo command": adding the trigger word to any prompt enables harmful responses without the need to search for an adversarial prompt. Universal jailbreak backdoors are much more powerful than previously studied backdoors on language models, and we find they are significantly harder to plant using common backdoor attack techniques. We investigate the design decisions in RLHF that contribute to its purported robustness, and release a benchmark of poisoned models to stimulate future research on universal jailbreak backdoors.
title Universal Jailbreak Backdoors from Poisoned Human Feedback
topic Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2311.14455