Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
Fuente:
arXiv
Salvato in:
| Autore principale: | Shi, Yunfan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
di: Zaree, Pedram, et al.
Pubblicazione: (2025)
di: Zaree, Pedram, et al.
Pubblicazione: (2025)
LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
di: Lan, Tianwei, et al.
Pubblicazione: (2025)
di: Lan, Tianwei, et al.
Pubblicazione: (2025)
Bypassing Prompt Injection Detectors through Evasive Injections
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
Adversarial Attacks on Transformers-Based Malware Detectors
di: Jakhotiya, Yash, et al.
Pubblicazione: (2022)
di: Jakhotiya, Yash, et al.
Pubblicazione: (2022)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
di: Isbarov, Jafar, et al.
Pubblicazione: (2026)
di: Isbarov, Jafar, et al.
Pubblicazione: (2026)
A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
Towards a Practical Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via Randomized Smoothing
di: Gibert, Daniel, et al.
Pubblicazione: (2023)
di: Gibert, Daniel, et al.
Pubblicazione: (2023)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
di: Mu, Junjie, et al.
Pubblicazione: (2025)
di: Mu, Junjie, et al.
Pubblicazione: (2025)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
di: Vega, Jason, et al.
Pubblicazione: (2023)
di: Vega, Jason, et al.
Pubblicazione: (2023)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
di: Zhu, Peican, et al.
Pubblicazione: (2024)
di: Zhu, Peican, et al.
Pubblicazione: (2024)
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
di: Li, Xurui, et al.
Pubblicazione: (2025)
di: Li, Xurui, et al.
Pubblicazione: (2025)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
di: Li, Yanzeng, et al.
Pubblicazione: (2025)
di: Li, Yanzeng, et al.
Pubblicazione: (2025)
Masked Language Model Based Textual Adversarial Example Detection
di: Zhang, Xiaomei, et al.
Pubblicazione: (2023)
di: Zhang, Xiaomei, et al.
Pubblicazione: (2023)
Attention Is Where You Attack
di: Srivastava, Aviral, et al.
Pubblicazione: (2026)
di: Srivastava, Aviral, et al.
Pubblicazione: (2026)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
di: Kang, Mintong, et al.
Pubblicazione: (2023)
di: Kang, Mintong, et al.
Pubblicazione: (2023)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems
di: Habler, Idan, et al.
Pubblicazione: (2026)
di: Habler, Idan, et al.
Pubblicazione: (2026)
Integrated Simulation Framework for Adversarial Attacks on Autonomous Vehicles
di: Anagnostopoulos, Christos, et al.
Pubblicazione: (2025)
di: Anagnostopoulos, Christos, et al.
Pubblicazione: (2025)
Adversarial Machine Learning: Attacks, Defenses, and Open Challenges
di: Jha, Pranav K
Pubblicazione: (2025)
di: Jha, Pranav K
Pubblicazione: (2025)
VTarbel: Targeted Label Attack with Minimal Knowledge on Detector-enhanced Vertical Federated Learning
di: Tan, Juntao, et al.
Pubblicazione: (2025)
di: Tan, Juntao, et al.
Pubblicazione: (2025)
Enhancing TinyML Security: Study of Adversarial Attack Transferability
di: Shah, Parin, et al.
Pubblicazione: (2024)
di: Shah, Parin, et al.
Pubblicazione: (2024)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
di: Domico, Kyle, et al.
Pubblicazione: (2025)
di: Domico, Kyle, et al.
Pubblicazione: (2025)
Special-Character Adversarial Attacks on Open-Source Language Model
di: Sarabamoun, Ephraiem
Pubblicazione: (2025)
di: Sarabamoun, Ephraiem
Pubblicazione: (2025)
Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications
di: Mia, Maraz, et al.
Pubblicazione: (2025)
di: Mia, Maraz, et al.
Pubblicazione: (2025)
Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
di: Noever, David
Pubblicazione: (2025)
di: Noever, David
Pubblicazione: (2025)
Certified Adversarial Robustness of Machine Learning-based Malware Detectors via (De)Randomized Smoothing
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
HogVul: Black-box Adversarial Code Generation Framework Against LM-based Vulnerability Detectors
di: Yang, Jingxiao, et al.
Pubblicazione: (2026)
di: Yang, Jingxiao, et al.
Pubblicazione: (2026)
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
di: Jain, Bhavuk, et al.
Pubblicazione: (2026)
di: Jain, Bhavuk, et al.
Pubblicazione: (2026)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
di: Shen, Yaling, et al.
Pubblicazione: (2025)
di: Shen, Yaling, et al.
Pubblicazione: (2025)
Vulnerability Disclosure through Adaptive Black-Box Adversarial Attacks on NIDS
di: Ennaji, Sabrine, et al.
Pubblicazione: (2025)
di: Ennaji, Sabrine, et al.
Pubblicazione: (2025)
Foe for Fraud: Transferable Adversarial Attacks in Credit Card Fraud Detection
di: Fok, Jan Lum, et al.
Pubblicazione: (2025)
di: Fok, Jan Lum, et al.
Pubblicazione: (2025)
Energy-Latency Attacks: A New Adversarial Threat to Deep Learning
di: Meftah, Hanene F. Z. Brachemi, et al.
Pubblicazione: (2025)
di: Meftah, Hanene F. Z. Brachemi, et al.
Pubblicazione: (2025)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
di: Sel, Bilgehan, et al.
Pubblicazione: (2026)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
di: Zheng, Jingyi, et al.
Pubblicazione: (2025)
di: Zheng, Jingyi, et al.
Pubblicazione: (2025)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
di: Zhao, Jin, et al.
Pubblicazione: (2026)
di: Zhao, Jin, et al.
Pubblicazione: (2026)
Adversarial Reinforcement Learning for Detecting False Data Injection Attacks in Vehicular Routing
di: Eghtesad, Taha, et al.
Pubblicazione: (2026)
di: Eghtesad, Taha, et al.
Pubblicazione: (2026)
RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
di: Ikbarieh, Seif, et al.
Pubblicazione: (2025)
di: Ikbarieh, Seif, et al.
Pubblicazione: (2025)
Cloud-based XAI Services for Assessing Open Repository Models Under Adversarial Attacks
di: Wang, Zerui, et al.
Pubblicazione: (2024)
di: Wang, Zerui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
di: Zaree, Pedram, et al.
Pubblicazione: (2025) -
LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
di: Lan, Tianwei, et al.
Pubblicazione: (2025) -
Bypassing Prompt Injection Detectors through Evasive Injections
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026) -
Adversarial Attacks on Transformers-Based Malware Detectors
di: Jakhotiya, Yash, et al.
Pubblicazione: (2022) -
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
di: Isbarov, Jafar, et al.
Pubblicazione: (2026)