Salvato in:
Dettagli Bibliografici
Autori principali: Xue, Anton, Khare, Avishree, Alur, Rajeev, Goel, Surbhi, Wong, Eric
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2407.00075
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917939780780032
author Xue, Anton
Khare, Avishree
Alur, Rajeev
Goel, Surbhi
Wong, Eric
author_facet Xue, Anton
Khare, Avishree
Alur, Rajeev
Goel, Surbhi
Wong, Eric
contents We study how to subvert large language models (LLMs) from following prompt-specified rules. We first formalize rule-following as inference in propositional Horn logic, a mathematical system in which rules have the form "if $P$ and $Q$, then $R$" for some propositions $P$, $Q$, and $R$. Next, we prove that although small transformers can faithfully follow such rules, maliciously crafted prompts can still mislead both theoretical constructions and models learned from data. Furthermore, we demonstrate that popular attack algorithms on LLMs find adversarial prompts and induce attention patterns that align with our theory. Our novel logic-based framework provides a foundation for studying LLMs in rule-based settings, enabling a formal analysis of tasks like logical reasoning and jailbreak attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00075
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
Xue, Anton
Khare, Avishree
Alur, Rajeev
Goel, Surbhi
Wong, Eric
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
We study how to subvert large language models (LLMs) from following prompt-specified rules. We first formalize rule-following as inference in propositional Horn logic, a mathematical system in which rules have the form "if $P$ and $Q$, then $R$" for some propositions $P$, $Q$, and $R$. Next, we prove that although small transformers can faithfully follow such rules, maliciously crafted prompts can still mislead both theoretical constructions and models learned from data. Furthermore, we demonstrate that popular attack algorithms on LLMs find adversarial prompts and induce attention patterns that align with our theory. Our novel logic-based framework provides a foundation for studying LLMs in rule-based settings, enabling a formal analysis of tasks like logical reasoning and jailbreak attacks.
title Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
topic Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2407.00075