Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Junyoung, Ju, Seongyong, Park, Sunghwan, Lee, Jaewoo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911734629924864
author Park, Junyoung
Ju, Seongyong
Park, Sunghwan
Lee, Jaewoo
author_facet Park, Junyoung
Ju, Seongyong
Park, Sunghwan
Lee, Jaewoo
contents As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety training. Traditional jailbreak techniques typically focus on a single prompt injection, neglecting the models' ability to remember the flow of conversation and the user's instructions. In this paper, we propose Persona Attack, a memory injection based jailbreak method that manipulates the model's context window through a step by step approach. Experimental results from applying Persona Attack to several widely used LLMs reveal that, as injections accumulate in memory, models increasingly prioritize these instructions over their internal safety alignment mechanisms. Furthermore, our experiments empirically demonstrate that the attack success rate varies not only according to the memory implementation of the model, but also combinations of instructions and can reach 95% under specific instruction configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00150
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
Park, Junyoung
Ju, Seongyong
Park, Sunghwan
Lee, Jaewoo
Cryptography and Security
Artificial Intelligence
As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety training. Traditional jailbreak techniques typically focus on a single prompt injection, neglecting the models' ability to remember the flow of conversation and the user's instructions. In this paper, we propose Persona Attack, a memory injection based jailbreak method that manipulates the model's context window through a step by step approach. Experimental results from applying Persona Attack to several widely used LLMs reveal that, as injections accumulate in memory, models increasingly prioritize these instructions over their internal safety alignment mechanisms. Furthermore, our experiments empirically demonstrate that the attack success rate varies not only according to the memory implementation of the model, but also combinations of instructions and can reach 95% under specific instruction configurations.
title Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2606.00150