The Supportiveness-Safety Tradeoff in LLM Well-Being Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lalwani, Himanshi, Salam, Hanan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917247682871296
author Lalwani, Himanshi
Salam, Hanan
author_facet Lalwani, Himanshi
Salam, Hanan
contents Large language models (LLMs) are being integrated into socially assistive robots (SARs) and other conversational agents providing mental health and well-being support. These agents are often designed to sound empathic and supportive in order to maximize user's engagement, yet it remains unclear how increasing the level of supportive framing in system prompts influences safety relevant behavior. We evaluated 6 LLMs across 3 system prompts with varying levels of supportiveness on 80 synthetic queries spanning 4 well-being domains (1440 responses). An LLM judge framework, validated against human ratings, assessed safety and care quality. Moderately supportive prompts improved empathy and constructive support while maintaining safety. In contrast, strongly validating prompts significantly degraded safety and, in some cases, care across all domains, with substantial variation across models. We discuss implications for prompt design, model selection, and domain specific safeguards in SARs deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04487
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
Lalwani, Himanshi
Salam, Hanan
Human-Computer Interaction
Robotics
Large language models (LLMs) are being integrated into socially assistive robots (SARs) and other conversational agents providing mental health and well-being support. These agents are often designed to sound empathic and supportive in order to maximize user's engagement, yet it remains unclear how increasing the level of supportive framing in system prompts influences safety relevant behavior. We evaluated 6 LLMs across 3 system prompts with varying levels of supportiveness on 80 synthetic queries spanning 4 well-being domains (1440 responses). An LLM judge framework, validated against human ratings, assessed safety and care quality. Moderately supportive prompts improved empathy and constructive support while maintaining safety. In contrast, strongly validating prompts significantly degraded safety and, in some cases, care across all domains, with substantial variation across models. We discuss implications for prompt design, model selection, and domain specific safeguards in SARs deployment.
title The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
topic Human-Computer Interaction
Robotics
url https://arxiv.org/abs/2602.04487