AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908732998287360 |
|---|---|
| author | Zhu, Sicheng Amos, Brandon Tian, Yuandong Guo, Chuan Evtimov, Ivan |
| author_facet | Zhu, Sicheng Amos, Brandon Tian, Yuandong Guo, Chuan Evtimov, Ivan |
| contents | Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforward, this objective has two limitations: limited control over model behaviors, yielding incomplete or unrealistic jailbroken responses, and a rigid format that hinders optimization. We introduce AdvPrefix, a plug-and-play prefix-forcing objective that selects one or more model-dependent prefixes by combining two criteria: high prefilling attack success rates and low negative log-likelihood. AdvPrefix integrates seamlessly into existing jailbreak attacks to mitigate the previous limitations for free. For example, replacing GCG's default prefixes on Llama-3 improves nuanced attack success rates from 14% to 80%, revealing that current safety alignment fails to generalize to new prefixes. Code and selected prefixes are released at github.com/facebookresearch/jailbreak-objectives. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_10321 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AdvPrefix: An Objective for Nuanced LLM Jailbreaks Zhu, Sicheng Amos, Brandon Tian, Yuandong Guo, Chuan Evtimov, Ivan Machine Learning Artificial Intelligence Computation and Language Cryptography and Security Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforward, this objective has two limitations: limited control over model behaviors, yielding incomplete or unrealistic jailbroken responses, and a rigid format that hinders optimization. We introduce AdvPrefix, a plug-and-play prefix-forcing objective that selects one or more model-dependent prefixes by combining two criteria: high prefilling attack success rates and low negative log-likelihood. AdvPrefix integrates seamlessly into existing jailbreak attacks to mitigate the previous limitations for free. For example, replacing GCG's default prefixes on Llama-3 improves nuanced attack success rates from 14% to 80%, revealing that current safety alignment fails to generalize to new prefixes. Code and selected prefixes are released at github.com/facebookresearch/jailbreak-objectives. |
| title | AdvPrefix: An Objective for Nuanced LLM Jailbreaks |
| topic | Machine Learning Artificial Intelligence Computation and Language Cryptography and Security |
| url | https://arxiv.org/abs/2412.10321 |