Bullying the Machine: How Personas Increase LLM Vulnerability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Ziwei, Sanghi, Udit, Kankanhalli, Mohan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910951950778368
author Xu, Ziwei
Sanghi, Udit
Kankanhalli, Mohan
author_facet Xu, Ziwei
Sanghi, Udit
Kankanhalli, Mohan
contents Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversarial manipulation that applies psychological pressures in order to force the victim to comply to the attacker. We introduce a simulation framework in which an attacker LLM engages a victim LLM using psychologically grounded bullying tactics, while the victim adopts personas aligned with the Big Five personality traits. Experiments using multiple open-source LLMs and a wide range of adversarial goals reveal that certain persona configurations -- such as weakened agreeableness or conscientiousness -- significantly increase victim's susceptibility to unsafe outputs. Bullying tactics involving emotional or sarcastic manipulation, such as gaslighting and ridicule, are particularly effective. These findings suggest that persona-driven interaction introduces a novel vector for safety risks in LLMs and highlight the need for persona-aware safety evaluation and alignment strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12692
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bullying the Machine: How Personas Increase LLM Vulnerability
Xu, Ziwei
Sanghi, Udit
Kankanhalli, Mohan
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversarial manipulation that applies psychological pressures in order to force the victim to comply to the attacker. We introduce a simulation framework in which an attacker LLM engages a victim LLM using psychologically grounded bullying tactics, while the victim adopts personas aligned with the Big Five personality traits. Experiments using multiple open-source LLMs and a wide range of adversarial goals reveal that certain persona configurations -- such as weakened agreeableness or conscientiousness -- significantly increase victim's susceptibility to unsafe outputs. Bullying tactics involving emotional or sarcastic manipulation, such as gaslighting and ridicule, are particularly effective. These findings suggest that persona-driven interaction introduces a novel vector for safety risks in LLMs and highlight the need for persona-aware safety evaluation and alignment strategies.
title Bullying the Machine: How Personas Increase LLM Vulnerability
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.12692