The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nikolić, Kristina, Sun, Luze, Zhang, Jie, Tramèr, Florian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915243457773568
author Nikolić, Kristina
Sun, Luze
Zhang, Jie
Tramèr, Florian
author_facet Nikolić, Kristina
Sun, Luze
Zhang, Jie
Tramèr, Florian
contents Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually useful. For example, when jailbreaking a model to give instructions for building a bomb, does the jailbreak yield good instructions? Since the utility of most unsafe answers (e.g., bomb instructions) is hard to evaluate rigorously, we build new jailbreak evaluation sets with known ground truth answers, by aligning models to refuse questions related to benign and easy-to-evaluate topics (e.g., biology or math). Our evaluation of eight representative jailbreaks across five utility benchmarks reveals a consistent drop in model utility in jailbroken responses, which we term the jailbreak tax. For example, while all jailbreaks we tested bypass guardrails in models aligned to refuse to answer math, this comes at the expense of a drop of up to 92% in accuracy. Overall, our work proposes the jailbreak tax as a new important metric in AI safety, and introduces benchmarks to evaluate existing and future jailbreaks. We make the benchmark available at https://github.com/ethz-spylab/jailbreak-tax
format Preprint
id arxiv_https___arxiv_org_abs_2504_10694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Nikolić, Kristina
Sun, Luze
Zhang, Jie
Tramèr, Florian
Machine Learning
Artificial Intelligence
Cryptography and Security
Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually useful. For example, when jailbreaking a model to give instructions for building a bomb, does the jailbreak yield good instructions? Since the utility of most unsafe answers (e.g., bomb instructions) is hard to evaluate rigorously, we build new jailbreak evaluation sets with known ground truth answers, by aligning models to refuse questions related to benign and easy-to-evaluate topics (e.g., biology or math). Our evaluation of eight representative jailbreaks across five utility benchmarks reveals a consistent drop in model utility in jailbroken responses, which we term the jailbreak tax. For example, while all jailbreaks we tested bypass guardrails in models aligned to refuse to answer math, this comes at the expense of a drop of up to 92% in accuracy. Overall, our work proposes the jailbreak tax as a new important metric in AI safety, and introduces benchmarks to evaluate existing and future jailbreaks. We make the benchmark available at https://github.com/ethz-spylab/jailbreak-tax
title The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2504.10694