When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jiahe, Guo, Xiangran, Hu, Yulin, Long, Zimo, Sui, Xingyu, Zhi, Xuda, Huang, Yongbo, He, Hao, Zhao, Weixiang, Zhao, Yanyan, Qin, Bing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691414962176
author Guo, Jiahe
Guo, Xiangran
Hu, Yulin
Long, Zimo
Sui, Xingyu
Zhi, Xuda
Huang, Yongbo
He, Hao
Zhao, Weixiang
Zhao, Yanyan
Qin, Bing
author_facet Guo, Jiahe
Guo, Xiangran
Hu, Yulin
Long, Zimo
Sui, Xingyu
Zhi, Xuda
Huang, Yongbo
He, Hao
Zhao, Weixiang
Zhao, Yanyan
Qin, Bing
contents Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized agents prioritizes utility and user experience, treating memory as a neutral component and largely overlooking its safety implications. In this paper, we reveal intent legitimation, a previously underexplored safety failure in personalized agents, where benign personal memories bias intent inference and cause models to legitimize inherently harmful queries. To study this phenomenon, we introduce PS-Bench, a benchmark designed to identify and quantify intent legitimation in personalized interactions. Across multiple memory-augmented agent frameworks and base LLMs, personalization increases attack success rates by 15.8\%--243.7\% relative to stateless baselines. We further provide mechanistic evidence for intent legitimation from internal representations space, and propose a lightweight detection-reflection method that effectively reduces safety degradation. Overall, our work provides the first systematic exploration and evaluation of intent legitimation as a safety failure mode that naturally arises from benign, real-world personalization, highlighting the importance of assessing safety under long-term personal context. Our code is available at: https://github.com/MuyuenLP/PS-Bench. WARNING: This paper may contain harmful content.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17887
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Guo, Jiahe
Guo, Xiangran
Hu, Yulin
Long, Zimo
Sui, Xingyu
Zhi, Xuda
Huang, Yongbo
He, Hao
Zhao, Weixiang
Zhao, Yanyan
Qin, Bing
Artificial Intelligence
Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized agents prioritizes utility and user experience, treating memory as a neutral component and largely overlooking its safety implications. In this paper, we reveal intent legitimation, a previously underexplored safety failure in personalized agents, where benign personal memories bias intent inference and cause models to legitimize inherently harmful queries. To study this phenomenon, we introduce PS-Bench, a benchmark designed to identify and quantify intent legitimation in personalized interactions. Across multiple memory-augmented agent frameworks and base LLMs, personalization increases attack success rates by 15.8\%--243.7\% relative to stateless baselines. We further provide mechanistic evidence for intent legitimation from internal representations space, and propose a lightweight detection-reflection method that effectively reduces safety degradation. Overall, our work provides the first systematic exploration and evaluation of intent legitimation as a safety failure mode that naturally arises from benign, real-world personalization, highlighting the importance of assessing safety under long-term personal context. Our code is available at: https://github.com/MuyuenLP/PS-Bench. WARNING: This paper may contain harmful content.
title When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2601.17887