Universal Jailbreak Suffixes Are Strong Attention Hijackers
Fuente:
arXiv
Saved in:
| Main Authors: | Ben-Tov, Matan, Geva, Mor, Sharif, Mahmood |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GASLITEing the Retrieval: Exploring Vulnerabilities in Dense Embedding-based Search
by: Ben-Tov, Matan, et al.
Published: (2024)
by: Ben-Tov, Matan, et al.
Published: (2024)
CaFA: Cost-aware, Feasible Attacks With Database Constraints Against Neural Tabular Classifiers
by: Ben-Tov, Matan, et al.
Published: (2025)
by: Ben-Tov, Matan, et al.
Published: (2025)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
Impactful Bit-Flip Search on Full-precision Models
by: Benedek, Nadav, et al.
Published: (2024)
by: Benedek, Nadav, et al.
Published: (2024)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
Inception Attacks: Immersive Hijacking in Virtual Reality Systems
by: Yang, Zhuolin, et al.
Published: (2024)
by: Yang, Zhuolin, et al.
Published: (2024)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
CAMH: Advancing Model Hijacking Attack in Machine Learning
by: He, Xing, et al.
Published: (2024)
by: He, Xing, et al.
Published: (2024)
Exploring Jamming and Hijacking Attacks for Micro Aerial Drones
by: Mekdad, Yassine, et al.
Published: (2024)
by: Mekdad, Yassine, et al.
Published: (2024)
Make Split, not Hijack: Preventing Feature-Space Hijacking Attacks in Split Learning
by: Khan, Tanveer, et al.
Published: (2024)
by: Khan, Tanveer, et al.
Published: (2024)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models
by: Liu, Yuansen, et al.
Published: (2026)
by: Liu, Yuansen, et al.
Published: (2026)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
by: Huang, Zhengxian, et al.
Published: (2026)
by: Huang, Zhengxian, et al.
Published: (2026)
Hijacking Attacks against Neural Networks by Analyzing Training Data
by: Ge, Yunjie, et al.
Published: (2024)
by: Ge, Yunjie, et al.
Published: (2024)
Towards Action Hijacking of Large Language Model-based Agent
by: Zhang, Yuyang, et al.
Published: (2024)
by: Zhang, Yuyang, et al.
Published: (2024)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor
by: Chang, Wenhan, et al.
Published: (2026)
by: Chang, Wenhan, et al.
Published: (2026)
A StrongREJECT for Empty Jailbreaks
by: Souly, Alexandra, et al.
Published: (2024)
by: Souly, Alexandra, et al.
Published: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Model Hijacking Attack in Federated Learning
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Vera Verto: Multimodal Hijacking Attack
by: Zhang, Minxing, et al.
Published: (2024)
by: Zhang, Minxing, et al.
Published: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning
by: He, Weiyang, et al.
Published: (2025)
by: He, Weiyang, et al.
Published: (2025)
Exploiting Sequence Number Leakage: TCP Hijacking in NAT-Enabled Wi-Fi Networks
by: Yang, Yuxiang, et al.
Published: (2024)
by: Yang, Yuxiang, et al.
Published: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Osmosis Distillation: Model Hijacking with the Fewest Samples
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
Resolving the Correct Library: A Loader-Level Defense Solution Against Shared Object Hijacking
by: Ozkan, Can, et al.
Published: (2026)
by: Ozkan, Can, et al.
Published: (2026)
EvilScreen Attack: Smart TV Hijacking via Multi-channel Remote Control Mimicry
by: Zhang, Yiwei, et al.
Published: (2022)
by: Zhang, Yiwei, et al.
Published: (2022)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
by: Zhang, Yucheng, et al.
Published: (2024)
by: Zhang, Yucheng, et al.
Published: (2024)
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
by: Wang, Ziwei, et al.
Published: (2026)
by: Wang, Ziwei, et al.
Published: (2026)
ClawGuard: Out-of-Band Detection of LLM Agent Workflow Hijacking via EM Side Channel
by: Gan, Leo Linqian, et al.
Published: (2026)
by: Gan, Leo Linqian, et al.
Published: (2026)
FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language Models
by: Yao, Dongyu, et al.
Published: (2023)
by: Yao, Dongyu, et al.
Published: (2023)
WHITE PAPER: A Brief Exploration of Data Exfiltration using GCG Suffixes
by: Valbuena, Victor
Published: (2024)
by: Valbuena, Victor
Published: (2024)
Moshi Moshi? A Model Selection Hijacking Adversarial Attack
by: Petrucci, Riccardo, et al.
Published: (2025)
by: Petrucci, Riccardo, et al.
Published: (2025)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
by: Zhao, Gejian, et al.
Published: (2025)
by: Zhao, Gejian, et al.
Published: (2025)
Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit
by: Sun, Qi, et al.
Published: (2026)
by: Sun, Qi, et al.
Published: (2026)
Similar Items
-
GASLITEing the Retrieval: Exploring Vulnerabilities in Dense Embedding-based Search
by: Ben-Tov, Matan, et al.
Published: (2024) -
CaFA: Cost-aware, Feasible Attacks With Database Constraints Against Neural Tabular Classifiers
by: Ben-Tov, Matan, et al.
Published: (2025) -
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026) -
Impactful Bit-Flip Search on Full-precision Models
by: Benedek, Nadav, et al.
Published: (2024) -
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)