Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Wenjie, Wang, Huandong, Gao, Junyao, Wan, Guoan, Jiang, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
von: Xin, Rui, et al.
Veröffentlicht: (2025)
von: Xin, Rui, et al.
Veröffentlicht: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions
von: Liu, Jinxin, et al.
Veröffentlicht: (2024)
von: Liu, Jinxin, et al.
Veröffentlicht: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Privacy Preserving In-Context-Learning Framework for Large Language Models
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
Analysis of Privacy Leakage in Federated Large Language Models
von: Vu, Minh N., et al.
Veröffentlicht: (2024)
von: Vu, Minh N., et al.
Veröffentlicht: (2024)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
Information Leakage from Embedding in Large Language Models
von: Wan, Zhipeng, et al.
Veröffentlicht: (2024)
von: Wan, Zhipeng, et al.
Veröffentlicht: (2024)
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
Topic-Based Watermarks for Large Language Models
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
von: Chua, Lynn, et al.
Veröffentlicht: (2024)
von: Chua, Lynn, et al.
Veröffentlicht: (2024)
A Probabilistic Fluctuation based Membership Inference Attack for Diffusion Models
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
von: Hughes, Anthony, et al.
Veröffentlicht: (2025)
von: Hughes, Anthony, et al.
Veröffentlicht: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Do Phone-Use Agents Respect Your Privacy?
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
On the Price of Privacy for Language Identification and Generation
von: Li, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2026)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
Model Provenance Testing for Large Language Models
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
How Private is Your Attention? Bridging Privacy with In-Context Learning
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Membership Inference Attacks and Privacy in Topic Modeling
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024) -
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023) -
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
von: Xin, Rui, et al.
Veröffentlicht: (2025) -
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026) -
Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions
von: Liu, Jinxin, et al.
Veröffentlicht: (2024)