Jailbreaking Large Language Models in Infinitely Many Ways

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goldstein, Oliver, La Malfa, Emanuele, Drinkall, Felix, Marro, Samuele, Wooldridge, Michael
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913734341951488
author Goldstein, Oliver
La Malfa, Emanuele
Drinkall, Felix
Marro, Samuele
Wooldridge, Michael
author_facet Goldstein, Oliver
La Malfa, Emanuele
Drinkall, Felix
Marro, Samuele
Wooldridge, Michael
contents We discuss the ``Infinitely Many Paraphrases'' attacks (IMP), a category of jailbreaks that leverages the increasing capabilities of a model to handle paraphrases and encoded communications to bypass their defensive mechanisms. IMPs' viability pairs and grows with a model's capabilities to handle and bind the semantics of simple mappings between tokens and work extremely well in practice, posing a concrete threat to the users of the most powerful LLMs in commerce. We show how one can bypass the safeguards of the most powerful open- and closed-source LLMs and generate content that explicitly violates their safety policies. One can protect against IMPs by improving the guardrails and making them scale with the LLMs' capabilities. For two categories of attacks that are straightforward to implement, i.e., bijection and encoding, we discuss two defensive strategies, one in token and the other in embedding space. We conclude with some research questions we believe should be prioritised to enhance the defensive mechanisms of LLMs and our understanding of their safety.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking Large Language Models in Infinitely Many Ways
Goldstein, Oliver
La Malfa, Emanuele
Drinkall, Felix
Marro, Samuele
Wooldridge, Michael
Machine Learning
Cryptography and Security
We discuss the ``Infinitely Many Paraphrases'' attacks (IMP), a category of jailbreaks that leverages the increasing capabilities of a model to handle paraphrases and encoded communications to bypass their defensive mechanisms. IMPs' viability pairs and grows with a model's capabilities to handle and bind the semantics of simple mappings between tokens and work extremely well in practice, posing a concrete threat to the users of the most powerful LLMs in commerce. We show how one can bypass the safeguards of the most powerful open- and closed-source LLMs and generate content that explicitly violates their safety policies. One can protect against IMPs by improving the guardrails and making them scale with the LLMs' capabilities. For two categories of attacks that are straightforward to implement, i.e., bijection and encoding, we discuss two defensive strategies, one in token and the other in embedding space. We conclude with some research questions we believe should be prioritised to enhance the defensive mechanisms of LLMs and our understanding of their safety.
title Jailbreaking Large Language Models in Infinitely Many Ways
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2501.10800