AdvPrefix: An Objective for Nuanced LLM Jailbreaks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Sicheng, Amos, Brandon, Tian, Yuandong, Guo, Chuan, Evtimov, Ivan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908732998287360
author Zhu, Sicheng
Amos, Brandon
Tian, Yuandong
Guo, Chuan
Evtimov, Ivan
author_facet Zhu, Sicheng
Amos, Brandon
Tian, Yuandong
Guo, Chuan
Evtimov, Ivan
contents Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforward, this objective has two limitations: limited control over model behaviors, yielding incomplete or unrealistic jailbroken responses, and a rigid format that hinders optimization. We introduce AdvPrefix, a plug-and-play prefix-forcing objective that selects one or more model-dependent prefixes by combining two criteria: high prefilling attack success rates and low negative log-likelihood. AdvPrefix integrates seamlessly into existing jailbreak attacks to mitigate the previous limitations for free. For example, replacing GCG's default prefixes on Llama-3 improves nuanced attack success rates from 14% to 80%, revealing that current safety alignment fails to generalize to new prefixes. Code and selected prefixes are released at github.com/facebookresearch/jailbreak-objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Zhu, Sicheng
Amos, Brandon
Tian, Yuandong
Guo, Chuan
Evtimov, Ivan
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforward, this objective has two limitations: limited control over model behaviors, yielding incomplete or unrealistic jailbroken responses, and a rigid format that hinders optimization. We introduce AdvPrefix, a plug-and-play prefix-forcing objective that selects one or more model-dependent prefixes by combining two criteria: high prefilling attack success rates and low negative log-likelihood. AdvPrefix integrates seamlessly into existing jailbreak attacks to mitigate the previous limitations for free. For example, replacing GCG's default prefixes on Llama-3 improves nuanced attack success rates from 14% to 80%, revealing that current safety alignment fails to generalize to new prefixes. Code and selected prefixes are released at github.com/facebookresearch/jailbreak-objectives.
title AdvPrefix: An Objective for Nuanced LLM Jailbreaks
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2412.10321