Attention Is Where You Attack
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Srivastava, Aviral, Panda, Sourav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation
von: Srivastava, Aviral, et al.
Veröffentlicht: (2024)
von: Srivastava, Aviral, et al.
Veröffentlicht: (2024)
When and Where do Data Poisons Attack Textual Inversion?
von: Styborski, Jeremy, et al.
Veröffentlicht: (2025)
von: Styborski, Jeremy, et al.
Veröffentlicht: (2025)
A Survey on Offensive AI Within Cybersecurity
von: Girhepuje, Sahil, et al.
Veröffentlicht: (2024)
von: Girhepuje, Sahil, et al.
Veröffentlicht: (2024)
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
von: Nakash, Itay, et al.
Veröffentlicht: (2024)
von: Nakash, Itay, et al.
Veröffentlicht: (2024)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
von: Shi, Yunfan
Veröffentlicht: (2024)
von: Shi, Yunfan
Veröffentlicht: (2024)
Not What You Asked For: Typographic Attacks in Household Robot Manipulation
von: Iranmanesh, Ali, et al.
Veröffentlicht: (2026)
von: Iranmanesh, Ali, et al.
Veröffentlicht: (2026)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
von: Jin, Haotian, et al.
Veröffentlicht: (2025)
von: Jin, Haotian, et al.
Veröffentlicht: (2025)
Temporal-Spatial Attention Network (TSAN) for DoS Attack Detection in Network Traffic
von: Kayode, Bisola Faith, et al.
Veröffentlicht: (2025)
von: Kayode, Bisola Faith, et al.
Veröffentlicht: (2025)
Confidence Is All You Need for MI Attacks
von: Sinha, Abhishek, et al.
Veröffentlicht: (2023)
von: Sinha, Abhishek, et al.
Veröffentlicht: (2023)
Buffer is All You Need: Defending Federated Learning against Backdoor Attacks under Non-iids via Buffering
von: Lyu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lyu, Xingyu, et al.
Veröffentlicht: (2025)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
The Surface You Test Is Not the Surface That Breaks
von: Arman, Shifat E, et al.
Veröffentlicht: (2026)
von: Arman, Shifat E, et al.
Veröffentlicht: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
von: Song, Baogang, et al.
Veröffentlicht: (2025)
von: Song, Baogang, et al.
Veröffentlicht: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
von: Kang, Mintong, et al.
Veröffentlicht: (2023)
von: Kang, Mintong, et al.
Veröffentlicht: (2023)
Attacking Slicing Network via Side-channel Reinforcement Learning Attack
von: Shao, Wei, et al.
Veröffentlicht: (2024)
von: Shao, Wei, et al.
Veröffentlicht: (2024)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
von: Pu, Rui, et al.
Veröffentlicht: (2024)
von: Pu, Rui, et al.
Veröffentlicht: (2024)
You Can Backdoor Personalized Federated Learning
von: Ye, Tiandi, et al.
Veröffentlicht: (2023)
von: Ye, Tiandi, et al.
Veröffentlicht: (2023)
Jailbreaking is (Mostly) Simpler Than You Think
von: Russinovich, Mark, et al.
Veröffentlicht: (2025)
von: Russinovich, Mark, et al.
Veröffentlicht: (2025)
Dynamic Target Attack
von: Xiu, Kedong, et al.
Veröffentlicht: (2025)
von: Xiu, Kedong, et al.
Veröffentlicht: (2025)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion
von: Xia, Shuhan, et al.
Veröffentlicht: (2025)
von: Xia, Shuhan, et al.
Veröffentlicht: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
I Don't Know You, But I Can Catch You: Real-Time Defense against Diverse Adversarial Patches for Object Detectors
von: Lin, Zijin, et al.
Veröffentlicht: (2024)
von: Lin, Zijin, et al.
Veröffentlicht: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
AttackER: Towards Enhancing Cyber-Attack Attribution with a Named Entity Recognition Dataset
von: Deka, Pritam, et al.
Veröffentlicht: (2024)
von: Deka, Pritam, et al.
Veröffentlicht: (2024)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
von: Liu, Zesen, et al.
Veröffentlicht: (2025)
von: Liu, Zesen, et al.
Veröffentlicht: (2025)
Where Do LLM-based Systems Break? A System-Level Security Framework for Risk Assessment and Treatment
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
von: Maple, Carsten, et al.
Veröffentlicht: (2026)
von: Maple, Carsten, et al.
Veröffentlicht: (2026)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
Security of Internet of Agents: Attacks and Countermeasures
von: Wang, Yuntao, et al.
Veröffentlicht: (2025)
von: Wang, Yuntao, et al.
Veröffentlicht: (2025)
Attacks on the neural network and defense methods
von: Korenev, A., et al.
Veröffentlicht: (2024)
von: Korenev, A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation
von: Srivastava, Aviral, et al.
Veröffentlicht: (2024) -
When and Where do Data Poisons Attack Textual Inversion?
von: Styborski, Jeremy, et al.
Veröffentlicht: (2025) -
A Survey on Offensive AI Within Cybersecurity
von: Girhepuje, Sahil, et al.
Veröffentlicht: (2024) -
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
von: Nakash, Itay, et al.
Veröffentlicht: (2024) -
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
von: Shi, Yunfan
Veröffentlicht: (2024)