No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Plaza-del-Arco, Flor Miriam, Röttger, Paul, Scherrer, Nino, Borgonovo, Emanuele, Plischke, Elmar, Hovy, Dirk
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914030623391744
author Plaza-del-Arco, Flor Miriam
Röttger, Paul
Scherrer, Nino
Borgonovo, Emanuele
Plischke, Elmar
Hovy, Dirk
author_facet Plaza-del-Arco, Flor Miriam
Röttger, Paul
Scherrer, Nino
Borgonovo, Emanuele
Plischke, Elmar
Hovy, Dirk
contents Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent work suggests that persona prompting can lead models to falsely refuse user requests. However, no work has fully quantified the extent of this issue. To address this gap, we measure the impact of 15 sociodemographic personas (based on gender, race, religion, and disability) on false refusal. To control for other factors, we also test 16 different models, 3 tasks (Natural Language Inference, politeness, and offensiveness classification), and nine prompt paraphrases. We propose a Monte Carlo-based method to quantify this issue in a sample-efficient manner. Our results show that as models become more capable, personas impact the refusal rate less and less. Certain sociodemographic personas increase false refusal in some models, which suggests underlying biases in the alignment strategies or safety mechanisms. However, we find that the model choice and task significantly influence false refusals, especially in sensitive content tasks. Our findings suggest that persona effects have been overestimated, and might be due to other factors.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
Plaza-del-Arco, Flor Miriam
Röttger, Paul
Scherrer, Nino
Borgonovo, Emanuele
Plischke, Elmar
Hovy, Dirk
Computation and Language
Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent work suggests that persona prompting can lead models to falsely refuse user requests. However, no work has fully quantified the extent of this issue. To address this gap, we measure the impact of 15 sociodemographic personas (based on gender, race, religion, and disability) on false refusal. To control for other factors, we also test 16 different models, 3 tasks (Natural Language Inference, politeness, and offensiveness classification), and nine prompt paraphrases. We propose a Monte Carlo-based method to quantify this issue in a sample-efficient manner. Our results show that as models become more capable, personas impact the refusal rate less and less. Certain sociodemographic personas increase false refusal in some models, which suggests underlying biases in the alignment strategies or safety mechanisms. However, we find that the model choice and task significantly influence false refusals, especially in sensitive content tasks. Our findings suggest that persona effects have been overestimated, and might be due to other factors.
title No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.08075