Privacy in Large Language Models: Attacks, Defenses and Future Directions
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoran, Chen, Yulin, Luo, Jinglong, Wang, Jiecong, Peng, Hao, Kang, Yan, Zhang, Xiaojin, Hu, Qi, Chan, Chunkit, Xu, Zenglin, Hooi, Bryan, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer Inference
by: Luo, Jinglong, et al.
Published: (2024)
by: Luo, Jinglong, et al.
Published: (2024)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Privacy-Preserved Neural Graph Databases
by: Hu, Qi, et al.
Published: (2023)
by: Hu, Qi, et al.
Published: (2023)
Activation-Guided Local Editing for Jailbreaking Attacks
by: Wang, Jiecong, et al.
Published: (2025)
by: Wang, Jiecong, et al.
Published: (2025)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
by: Fan, Wei, et al.
Published: (2024)
by: Fan, Wei, et al.
Published: (2024)
Automated Phishing Detection Using URLs and Webpages
by: Wang, Huilin, et al.
Published: (2024)
by: Wang, Huilin, et al.
Published: (2024)
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Privacy-Preserving Large Language Models: Mechanisms, Applications, and Future Directions
by: Zhao, Guoshenghui, et al.
Published: (2024)
by: Zhao, Guoshenghui, et al.
Published: (2024)
Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions
by: Xu, Yuming, et al.
Published: (2026)
by: Xu, Yuming, et al.
Published: (2026)
Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses
by: Song, Minkyoo, et al.
Published: (2026)
by: Song, Minkyoo, et al.
Published: (2026)
SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC
by: Luo, Jinglong, et al.
Published: (2024)
by: Luo, Jinglong, et al.
Published: (2024)
Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning Attack
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
On the Security and Privacy of Federated Learning: A Survey with Attacks, Defenses, Frameworks, Applications, and Future Directions
by: Jimenez-Gutierrez, Daniel M., et al.
Published: (2025)
by: Jimenez-Gutierrez, Daniel M., et al.
Published: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
Optimal Attack and Defense for Reinforcement Learning
by: McMahan, Jeremy, et al.
Published: (2023)
by: McMahan, Jeremy, et al.
Published: (2023)
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
by: Liu, Deng, et al.
Published: (2026)
by: Liu, Deng, et al.
Published: (2026)
Fed-AugMix: Balancing Privacy and Utility via Data Augmentation
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
Deciphering the Interplay between Attack and Protection Complexity in Privacy-Preserving Federated Learning
by: Zhang, Xiaojin, et al.
Published: (2025)
by: Zhang, Xiaojin, et al.
Published: (2025)
Detection and Defense Against Prominent Attacks on Preconditioned LLM-Integrated Virtual Assistants
by: Chan, Chun Fai, et al.
Published: (2024)
by: Chan, Chun Fai, et al.
Published: (2024)
System Prompt Extraction Attacks and Defenses in Large Language Models
by: Das, Badhan Chandra, et al.
Published: (2025)
by: Das, Badhan Chandra, et al.
Published: (2025)
JNI Global References Are Still Vulnerable: Attacks and Defenses
by: He, Yi, et al.
Published: (2024)
by: He, Yi, et al.
Published: (2024)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
Dual Defense: Enhancing Privacy and Mitigating Poisoning Attacks in Federated Learning
by: Xu, Runhua, et al.
Published: (2025)
by: Xu, Runhua, et al.
Published: (2025)
Multimodal Large Language Models for Phishing Webpage Detection and Identification
by: Lee, Jehyun, et al.
Published: (2024)
by: Lee, Jehyun, et al.
Published: (2024)
Advancing Differential Privacy: Where We Are Now and Future Directions for Real-World Deployment
by: Cummings, Rachel, et al.
Published: (2023)
by: Cummings, Rachel, et al.
Published: (2023)
Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
by: Xue, Jing, et al.
Published: (2025)
by: Xue, Jing, et al.
Published: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
SUAD: Solid-Channel Ultrasound Injection Attack and Defense to Voice Assistants
by: Liu, Chao, et al.
Published: (2025)
by: Liu, Chao, et al.
Published: (2025)
Sanitization of Multimedia Content: A Survey of Techniques, Attacks, and Future Directions
by: Ciccotelli, Andrea, et al.
Published: (2022)
by: Ciccotelli, Andrea, et al.
Published: (2022)
Similar Items
-
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
by: Chen, Yulin, et al.
Published: (2025) -
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
by: Chen, Yulin, et al.
Published: (2024) -
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024) -
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
by: Chen, Yulin, et al.
Published: (2025) -
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)