Defensive Dual Masking for Robust Adversarial Defense
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Wangli, Yang, Jie, Guo, Yi, Barthelemy, Johan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Attacks and Defense for Conversation Entailment Task
von: Yang, Zhenning, et al.
Veröffentlicht: (2024)
von: Yang, Zhenning, et al.
Veröffentlicht: (2024)
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models
von: Liang, Yuhan, et al.
Veröffentlicht: (2024)
von: Liang, Yuhan, et al.
Veröffentlicht: (2024)
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026)
von: Wen, Yule, et al.
Veröffentlicht: (2026)
A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLM
von: Wang, Bo, et al.
Veröffentlicht: (2026)
von: Wang, Bo, et al.
Veröffentlicht: (2026)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
von: Patel, Het, et al.
Veröffentlicht: (2025)
von: Patel, Het, et al.
Veröffentlicht: (2025)
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers
von: Peng, Zuquan, et al.
Veröffentlicht: (2024)
von: Peng, Zuquan, et al.
Veröffentlicht: (2024)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
von: Kim, Minkyoung, et al.
Veröffentlicht: (2024)
von: Kim, Minkyoung, et al.
Veröffentlicht: (2024)
Provable Defense Framework for LLM Jailbreaks via Noise-Augumented Alignment
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
Multilingual Collaborative Defense for Large Language Models
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
von: Sabir, Bushra, et al.
Veröffentlicht: (2023)
von: Sabir, Bushra, et al.
Veröffentlicht: (2023)
CertMask: Certifiable Defense Against Adversarial Patches via Theoretically Optimal Mask Coverage
von: Lyu, Xuntao, et al.
Veröffentlicht: (2025)
von: Lyu, Xuntao, et al.
Veröffentlicht: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
von: Ouyang, Yang, et al.
Veröffentlicht: (2025)
von: Ouyang, Yang, et al.
Veröffentlicht: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
von: Dong, Zhichen, et al.
Veröffentlicht: (2024)
von: Dong, Zhichen, et al.
Veröffentlicht: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks
von: Sivaroopan, Nirhoshan, et al.
Veröffentlicht: (2026)
von: Sivaroopan, Nirhoshan, et al.
Veröffentlicht: (2026)
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
von: Yang, April, et al.
Veröffentlicht: (2024)
von: Yang, April, et al.
Veröffentlicht: (2024)
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained Encoders
von: Lei, Zhuxin, et al.
Veröffentlicht: (2026)
von: Lei, Zhuxin, et al.
Veröffentlicht: (2026)
Defensive M2S: Training Guardrail Models on Compressed Multi-turn Conversations
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
von: Volkova, Svitlana, et al.
Veröffentlicht: (2025)
von: Volkova, Svitlana, et al.
Veröffentlicht: (2025)
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models
von: Kim, San, et al.
Veröffentlicht: (2026)
von: Kim, San, et al.
Veröffentlicht: (2026)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
von: Wang, Wenxiao, et al.
Veröffentlicht: (2025)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
A Survey on Agentic Security: Applications, Threats and Defenses
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
von: Piet, Julien, et al.
Veröffentlicht: (2023)
von: Piet, Julien, et al.
Veröffentlicht: (2023)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
Pruning as a Defense: Reducing Memorization in Large Language Models
von: Gupta, Mansi, et al.
Veröffentlicht: (2025)
von: Gupta, Mansi, et al.
Veröffentlicht: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adversarial Attacks and Defense for Conversation Entailment Task
von: Yang, Zhenning, et al.
Veröffentlicht: (2024) -
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models
von: Liang, Yuhan, et al.
Veröffentlicht: (2024) -
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026) -
A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLM
von: Wang, Bo, et al.
Veröffentlicht: (2026) -
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
von: Patel, Het, et al.
Veröffentlicht: (2025)