FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kuznetsov, Daniel, Cohen, Ofir, Shistik, Karin, Puzis, Rami, Shabtai, Asaf
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911571173703680
author Kuznetsov, Daniel
Cohen, Ofir
Shistik, Karin
Puzis, Rami
Shabtai, Asaf
author_facet Kuznetsov, Daniel
Cohen, Ofir
Shistik, Karin
Puzis, Rami
Shabtai, Asaf
contents Safety-aligned LLMs go through refusal training to reject harmful requests, but whether these mechanisms remain effective under emotionally charged stimuli is unexplored. We introduce FreakOut-LLM, a framework investigating whether emotional context compromises safety alignment in adversarial settings. Using validated psychological stimuli, we evaluate how emotional priming through system prompts affects jailbreak susceptibility across ten LLMs. We test three conditions (stress, relaxation, neutral) using scenarios from established psychological protocols, plus a no-prompt baseline, and evaluate attack success using HarmBench on AdvBench prompts. Stress priming increases jailbreak success by 65.2\% compared to neutral conditions (z = 5.93, p < 0.001; OR = 1.67, Cohen's d = 0.28), while relaxation priming produces no effect (p = 0.84). Five of ten models show significant vulnerability, with the largest effects concentrated in open-weight models. Logistic regression on 59,800 queries confirms stress as the sole significant condition predictor after controlling for prompt length (p = 0.61) and model identity. Measured psychological state strongly predicts attack success (|r|\geq0.70 across five instruments; all p < 0.001 in individual-level logistic regression). These results establish emotional context as a measurable attack surface with implications for real-world AI deployment in high-stress domains.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04992
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
Kuznetsov, Daniel
Cohen, Ofir
Shistik, Karin
Puzis, Rami
Shabtai, Asaf
Cryptography and Security
Artificial Intelligence
Safety-aligned LLMs go through refusal training to reject harmful requests, but whether these mechanisms remain effective under emotionally charged stimuli is unexplored. We introduce FreakOut-LLM, a framework investigating whether emotional context compromises safety alignment in adversarial settings. Using validated psychological stimuli, we evaluate how emotional priming through system prompts affects jailbreak susceptibility across ten LLMs. We test three conditions (stress, relaxation, neutral) using scenarios from established psychological protocols, plus a no-prompt baseline, and evaluate attack success using HarmBench on AdvBench prompts. Stress priming increases jailbreak success by 65.2\% compared to neutral conditions (z = 5.93, p < 0.001; OR = 1.67, Cohen's d = 0.28), while relaxation priming produces no effect (p = 0.84). Five of ten models show significant vulnerability, with the largest effects concentrated in open-weight models. Logistic regression on 59,800 queries confirms stress as the sole significant condition predictor after controlling for prompt length (p = 0.61) and model identity. Measured psychological state strongly predicts attack success (|r|\geq0.70 across five instruments; all p < 0.001 in individual-level logistic regression). These results establish emotional context as a measurable attack surface with implications for real-world AI deployment in high-stress domains.
title FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2604.04992