BadGPT-4o: stripping safety finetuning from GPT models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Krupkina, Ekaterina, Volkov, Dmitrii
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909419119312896
author Krupkina, Ekaterina
Volkov, Dmitrii
author_facet Krupkina, Ekaterina
Volkov, Dmitrii
contents We show a version of Qi et al. 2023's simple fine-tuning poisoning technique strips GPT-4o's safety guardrails without degrading the model. The BadGPT attack matches best white-box jailbreaks on HarmBench and StrongREJECT. It suffers no token overhead or performance hits common to jailbreaks, as evaluated on tinyMMLU and open-ended generations. Despite having been known for a year, this attack remains easy to execute.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05346
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BadGPT-4o: stripping safety finetuning from GPT models
Krupkina, Ekaterina
Volkov, Dmitrii
Cryptography and Security
Machine Learning
We show a version of Qi et al. 2023's simple fine-tuning poisoning technique strips GPT-4o's safety guardrails without degrading the model. The BadGPT attack matches best white-box jailbreaks on HarmBench and StrongREJECT. It suffers no token overhead or performance hits common to jailbreaks, as evaluated on tinyMMLU and open-ended generations. Despite having been known for a year, this attack remains easy to execute.
title BadGPT-4o: stripping safety finetuning from GPT models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2412.05346