Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gibbs, Tom, Kosak-Hine, Ethan, Ingebretsen, George, Zhang, Jason, Broomfield, Julius, Pieri, Sara, Iranmanesh, Reihaneh, Rabbany, Reihaneh, Pelrine, Kellin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914931684671488
author Gibbs, Tom
Kosak-Hine, Ethan
Ingebretsen, George
Zhang, Jason
Broomfield, Julius
Pieri, Sara
Iranmanesh, Reihaneh
Rabbany, Reihaneh
Pelrine, Kellin
author_facet Gibbs, Tom
Kosak-Hine, Ethan
Ingebretsen, George
Zhang, Jason
Broomfield, Julius
Pieri, Sara
Iranmanesh, Reihaneh
Rabbany, Reihaneh
Pelrine, Kellin
contents Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00137
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
Gibbs, Tom
Kosak-Hine, Ethan
Ingebretsen, George
Zhang, Jason
Broomfield, Julius
Pieri, Sara
Iranmanesh, Reihaneh
Rabbany, Reihaneh
Pelrine, Kellin
Cryptography and Security
Artificial Intelligence
Computation and Language
Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so.
title Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.00137