Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914931684671488 |
|---|---|
| author | Gibbs, Tom Kosak-Hine, Ethan Ingebretsen, George Zhang, Jason Broomfield, Julius Pieri, Sara Iranmanesh, Reihaneh Rabbany, Reihaneh Pelrine, Kellin |
| author_facet | Gibbs, Tom Kosak-Hine, Ethan Ingebretsen, George Zhang, Jason Broomfield, Julius Pieri, Sara Iranmanesh, Reihaneh Rabbany, Reihaneh Pelrine, Kellin |
| contents | Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_00137 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks Gibbs, Tom Kosak-Hine, Ethan Ingebretsen, George Zhang, Jason Broomfield, Julius Pieri, Sara Iranmanesh, Reihaneh Rabbany, Reihaneh Pelrine, Kellin Cryptography and Security Artificial Intelligence Computation and Language Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so. |
| title | Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks |
| topic | Cryptography and Security Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2409.00137 |