Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zheng, Mingqian, Hu, Wenjia, Zhao, Patrick, Eslami, Motahhare, Hwang, Jena D., Brahman, Faeze, Rose, Carolyn, Sap, Maarten
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908689243308032
author Zheng, Mingqian
Hu, Wenjia
Zhao, Patrick
Eslami, Motahhare
Hwang, Jena D.
Brahman, Faeze
Rose, Carolyn
Sap, Maarten
author_facet Zheng, Mingqian
Hu, Wenjia
Zhao, Patrick
Eslami, Motahhare
Hwang, Jena D.
Brahman, Faeze
Rose, Carolyn
Sap, Maarten
contents Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840 query-response pairs, we examine how different refusal strategies affect user perceptions across varying motivations. Our findings reveal that response strategy largely shapes user experience, while actual user motivation has negligible impact. Partial compliance -- providing general information without actionable details -- emerges as the optimal strategy, reducing negative user perceptions by over 50% to flat-out refusals. Complementing this, we analyze response patterns of 9 state-of-the-art LLMs and evaluate how 6 reward models score different refusal strategies, demonstrating that models rarely deploy partial compliance naturally and reward models currently undervalue it. This work demonstrates that effective guardrails require focusing on crafting thoughtful refusals rather than detecting intent, offering a path toward AI safety mechanisms that ensure both safety and sustained user engagement.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
Zheng, Mingqian
Hu, Wenjia
Zhao, Patrick
Eslami, Motahhare
Hwang, Jena D.
Brahman, Faeze
Rose, Carolyn
Sap, Maarten
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840 query-response pairs, we examine how different refusal strategies affect user perceptions across varying motivations. Our findings reveal that response strategy largely shapes user experience, while actual user motivation has negligible impact. Partial compliance -- providing general information without actionable details -- emerges as the optimal strategy, reducing negative user perceptions by over 50% to flat-out refusals. Complementing this, we analyze response patterns of 9 state-of-the-art LLMs and evaluate how 6 reward models score different refusal strategies, demonstrating that models rarely deploy partial compliance naturally and reward models currently undervalue it. This work demonstrates that effective guardrails require focusing on crafting thoughtful refusals rather than detecting intent, offering a path toward AI safety mechanisms that ensure both safety and sustained user engagement.
title Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2506.00195