When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Jiahe, Guo, Xiangran, Hu, Yulin, Long, Zimo, Sui, Xingyu, Zhi, Xuda, Huang, Yongbo, He, Hao, Zhao, Weixiang, Zhao, Yanyan, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
by: Hu, Yulin, et al.
Published: (2026)
by: Hu, Yulin, et al.
Published: (2026)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026)
by: Sui, Xingyu, et al.
Published: (2026)
Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
On Safety Risks in Experience-Driven Self-Evolving Agents
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
by: Guo, Jiahe, et al.
Published: (2026)
by: Guo, Jiahe, et al.
Published: (2026)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
by: Zhao, Weixiang, et al.
Published: (2024)
by: Zhao, Weixiang, et al.
Published: (2024)
Lens: Rethinking Multilingual Enhancement for Large Language Models
by: Zhao, Weixiang, et al.
Published: (2024)
by: Zhao, Weixiang, et al.
Published: (2024)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporter
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Learning to Learn from Multimodal Experience
by: Sui, Xingyu, et al.
Published: (2026)
by: Sui, Xingyu, et al.
Published: (2026)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
by: Fu, Xing, et al.
Published: (2026)
by: Fu, Xing, et al.
Published: (2026)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues
by: Sun, Lei, et al.
Published: (2024)
by: Sun, Lei, et al.
Published: (2024)
Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
by: Su, Miao, et al.
Published: (2026)
by: Su, Miao, et al.
Published: (2026)
Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence
by: Zhao, Weixiang, et al.
Published: (2024)
by: Zhao, Weixiang, et al.
Published: (2024)
RKLD: Reverse KL-Divergence-based Knowledge Distillation for Unlearning Personal Information in Large Language Models
by: Wang, Bichen, et al.
Published: (2024)
by: Wang, Bichen, et al.
Published: (2024)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language Models
by: Zhao, Weixiang, et al.
Published: (2024)
by: Zhao, Weixiang, et al.
Published: (2024)
Text-Driven Emotionally Continuous Talking Face Generation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Rethinking Experience Utilization in Self-Evolving Language Model Agents
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Semantics Over Syntax: Uncovering Pre-Authentication 5G Baseband Vulnerabilities
by: Huang, Qiqing, et al.
Published: (2026)
by: Huang, Qiqing, et al.
Published: (2026)
SafeScreen: A Safety-First Screening Framework for Personalized Video Retrieval for Vulnerable Users
by: Zhao, Wenzheng, et al.
Published: (2026)
by: Zhao, Wenzheng, et al.
Published: (2026)
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
by: Cheng, Sijia, et al.
Published: (2025)
by: Cheng, Sijia, et al.
Published: (2025)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
by: Wu, Yaozu, et al.
Published: (2025)
by: Wu, Yaozu, et al.
Published: (2025)
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents
by: Tan, Zhen, et al.
Published: (2025)
by: Tan, Zhen, et al.
Published: (2025)
Fragile by Design: On the Limits of Adversarial Defenses in Personalized Generation
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Enhancing Personality Recognition in Dialogue by Data Augmentation and Heterogeneous Conversational Graph Networks
by: Fu, Yahui, et al.
Published: (2024)
by: Fu, Yahui, et al.
Published: (2024)
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
Reinforcement Learning for Personalized Dialogue Management
by: Hengst, Floris den, et al.
Published: (2019)
by: Hengst, Floris den, et al.
Published: (2019)
Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
by: Huang, Zhaopei, et al.
Published: (2025)
by: Huang, Zhaopei, et al.
Published: (2025)
Similar Items
-
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
by: Hu, Yulin, et al.
Published: (2026) -
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026) -
Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
by: Zhao, Weixiang, et al.
Published: (2025) -
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025) -
On Safety Risks in Experience-Driven Self-Evolving Agents
by: Zhao, Weixiang, et al.
Published: (2026)