Attack and defense techniques in large language models: A survey and new perspectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Zhiyu, Chen, Kang, Lin, Yuanguo, Li, Kangkang, Liu, Yunxuan, Chen, Hefeng, Huang, Xingwang, Yu, Yuanhui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Data Security in Large Language Models
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
Optimizing watermarks for large language models
von: Wouters, Bram
Veröffentlicht: (2023)
von: Wouters, Bram
Veröffentlicht: (2023)
Proactive defense against LLM Jailbreak
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
von: Yi, Xin, et al.
Veröffentlicht: (2025)
von: Yi, Xin, et al.
Veröffentlicht: (2025)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
von: Xian, R. Patrick, et al.
Veröffentlicht: (2024)
von: Xian, R. Patrick, et al.
Veröffentlicht: (2024)
A Survey on Privacy Risks and Protection in Large Language Models
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
von: Ghosh, Rikhiya, et al.
Veröffentlicht: (2024)
von: Ghosh, Rikhiya, et al.
Veröffentlicht: (2024)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
Can a large language model be a gaslighter?
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
FORAY: Towards Effective Attack Synthesis against Deep Logical Vulnerabilities in DeFi Protocols
von: Wen, Hongbo, et al.
Veröffentlicht: (2024)
von: Wen, Hongbo, et al.
Veröffentlicht: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
von: Upadhayay, Bibek, et al.
Veröffentlicht: (2024)
von: Upadhayay, Bibek, et al.
Veröffentlicht: (2024)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
von: Afane, Mohamed, et al.
Veröffentlicht: (2025)
von: Afane, Mohamed, et al.
Veröffentlicht: (2025)
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
von: Chen, Yining, et al.
Veröffentlicht: (2026)
von: Chen, Yining, et al.
Veröffentlicht: (2026)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks
von: Chen, Yiyi, et al.
Veröffentlicht: (2024)
von: Chen, Yiyi, et al.
Veröffentlicht: (2024)
Reverse-Engineering Model Editing on Language Models
von: Sun, Zhiyu, et al.
Veröffentlicht: (2026)
von: Sun, Zhiyu, et al.
Veröffentlicht: (2026)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
von: Chen, Yunhao, et al.
Veröffentlicht: (2025)
Enhance Robustness of Language Models Against Variation Attack through Graph Integration
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
MPMA: Preference Manipulation Attack Against Model Context Protocol
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
Attacks on the neural network and defense methods
von: Korenev, A., et al.
Veröffentlicht: (2024)
von: Korenev, A., et al.
Veröffentlicht: (2024)
BadEdit: Backdooring large language models by model editing
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
von: Li, Yanzhou, et al.
Veröffentlicht: (2024)
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
von: Chuang, Ko-Wei, et al.
Veröffentlicht: (2025)
von: Chuang, Ko-Wei, et al.
Veröffentlicht: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Survey on Data Security in Large Language Models
von: Chen, Kang, et al.
Veröffentlicht: (2025) -
Optimizing watermarks for large language models
von: Wouters, Bram
Veröffentlicht: (2023) -
Proactive defense against LLM Jailbreak
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025) -
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
von: Yi, Xin, et al.
Veröffentlicht: (2025) -
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
von: Xian, R. Patrick, et al.
Veröffentlicht: (2024)