Prompts have evil twins

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Melamed, Rimon, McCabe, Lucas H., Wakhare, Tanay, Kim, Yejin, Huang, H. Howie, Boix-Adsera, Enric
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909337853624320
author Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
author_facet Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
contents We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" because they are obfuscated and uninterpretable (evil), but at the same time mimic the functionality of the original natural-language prompts (twins). Remarkably, evil twins transfer between models. We find these prompts by solving a maximum-likelihood problem which has applications of independent interest.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07064
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Prompts have evil twins
Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
Computation and Language
We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" because they are obfuscated and uninterpretable (evil), but at the same time mimic the functionality of the original natural-language prompts (twins). Remarkably, evil twins transfer between models. We find these prompts by solving a maximum-likelihood problem which has applications of independent interest.
title Prompts have evil twins
topic Computation and Language
url https://arxiv.org/abs/2311.07064