Removing RLHF Protections in GPT-4 via Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Qiusi, Fang, Richard, Bindu, Rohan, Gupta, Akul, Hashimoto, Tatsunori, Kang, Daniel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916195723116544
author Zhan, Qiusi
Fang, Richard
Bindu, Rohan
Gupta, Akul
Hashimoto, Tatsunori
Kang, Daniel
author_facet Zhan, Qiusi
Fang, Richard
Bindu, Rohan
Gupta, Akul
Hashimoto, Tatsunori
Kang, Daniel
contents As large language models (LLMs) have increased in their capabilities, so does their potential for dual use. To reduce harmful outputs, produces and vendors of LLMs have used reinforcement learning with human feedback (RLHF). In tandem, LLM vendors have been increasingly enabling fine-tuning of their most powerful models. However, concurrent work has shown that fine-tuning can remove RLHF protections. We may expect that the most powerful models currently available (GPT-4) are less susceptible to fine-tuning attacks. In this work, we show the contrary: fine-tuning allows attackers to remove RLHF protections with as few as 340 examples and a 95% success rate. These training examples can be automatically generated with weaker models. We further show that removing RLHF protections does not decrease usefulness on non-censored outputs, providing evidence that our fine-tuning strategy does not decrease usefulness despite using weaker models to generate training data. Our results show the need for further research on protections on LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_05553
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Removing RLHF Protections in GPT-4 via Fine-Tuning
Zhan, Qiusi
Fang, Richard
Bindu, Rohan
Gupta, Akul
Hashimoto, Tatsunori
Kang, Daniel
Computation and Language
Artificial Intelligence
As large language models (LLMs) have increased in their capabilities, so does their potential for dual use. To reduce harmful outputs, produces and vendors of LLMs have used reinforcement learning with human feedback (RLHF). In tandem, LLM vendors have been increasingly enabling fine-tuning of their most powerful models. However, concurrent work has shown that fine-tuning can remove RLHF protections. We may expect that the most powerful models currently available (GPT-4) are less susceptible to fine-tuning attacks. In this work, we show the contrary: fine-tuning allows attackers to remove RLHF protections with as few as 340 examples and a 95% success rate. These training examples can be automatically generated with weaker models. We further show that removing RLHF protections does not decrease usefulness on non-censored outputs, providing evidence that our fine-tuning strategy does not decrease usefulness despite using weaker models to generate training data. Our results show the need for further research on protections on LLMs.
title Removing RLHF Protections in GPT-4 via Fine-Tuning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.05553