Prompts have evil twins

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Melamed, Rimon, McCabe, Lucas H., Wakhare, Tanay, Kim, Yejin, Huang, H. Howie, Boix-Adsera, Enric
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909337853624320
author Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
author_facet Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
contents We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" because they are obfuscated and uninterpretable (evil), but at the same time mimic the functionality of the original natural-language prompts (twins). Remarkably, evil twins transfer between models. We find these prompts by solving a maximum-likelihood problem which has applications of independent interest.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07064
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Prompts have evil twins
Melamed, Rimon
McCabe, Lucas H.
Wakhare, Tanay
Kim, Yejin
Huang, H. Howie
Boix-Adsera, Enric
Computation and Language
We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" because they are obfuscated and uninterpretable (evil), but at the same time mimic the functionality of the original natural-language prompts (twins). Remarkably, evil twins transfer between models. We find these prompts by solving a maximum-likelihood problem which has applications of independent interest.
title Prompts have evil twins
topic Computation and Language
url https://arxiv.org/abs/2311.07064