Contextualized Privacy Defense for LLM Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wen, Yule, Zhang, Yanzhe, Lian, Jianxun, Yi, Xiaoyuan, Xie, Xing, Yang, Diyi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914365487185920
author Wen, Yule
Zhang, Yanzhe
Lian, Jianxun
Yi, Xiaoyuan
Xie, Xing
Yang, Diyi
author_facet Wen, Yule
Zhang, Yanzhe
Lian, Jianxun
Yi, Xiaoyuan
Xie, Xing
Yang, Diyi
contents LLM agents increasingly act on users' personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and guarding. These paradigms are insufficient for supporting contextual, proactive privacy decisions in multi-step agent execution. We propose Contextualized Defense Instructing (CDI), a new privacy defense paradigm in which an instructor model generates step-specific, context-aware privacy guidance during execution, proactively shaping actions rather than merely constraining or vetoing them. Crucially, CDI is paired with an experience-driven optimization framework that trains the instructor via reinforcement learning (RL), where we convert failure trajectories with privacy violations into learning environments. We formalize baseline defenses and CDI as distinct intervention points in a canonical agent loop, and compare their privacy-helpfulness trade-offs within a unified simulation framework. Results show that our CDI consistently achieves a better balance between privacy preservation (94.2%) and helpfulness (80.6%) than baselines, with superior robustness to adversarial conditions and generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02983
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contextualized Privacy Defense for LLM Agents
Wen, Yule
Zhang, Yanzhe
Lian, Jianxun
Yi, Xiaoyuan
Xie, Xing
Yang, Diyi
Cryptography and Security
Artificial Intelligence
Computation and Language
LLM agents increasingly act on users' personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and guarding. These paradigms are insufficient for supporting contextual, proactive privacy decisions in multi-step agent execution. We propose Contextualized Defense Instructing (CDI), a new privacy defense paradigm in which an instructor model generates step-specific, context-aware privacy guidance during execution, proactively shaping actions rather than merely constraining or vetoing them. Crucially, CDI is paired with an experience-driven optimization framework that trains the instructor via reinforcement learning (RL), where we convert failure trajectories with privacy violations into learning environments. We formalize baseline defenses and CDI as distinct intervention points in a canonical agent loop, and compare their privacy-helpfulness trade-offs within a unified simulation framework. Results show that our CDI consistently achieves a better balance between privacy preservation (94.2%) and helpfulness (80.6%) than baselines, with superior robustness to adversarial conditions and generalization.
title Contextualized Privacy Defense for LLM Agents
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.02983