Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ziwei, Chen, Jing, Liang, Ruichao, Wang, Zhi, Feng, Yebo, Jia, Ju, Du, Ruiying, Wu, Cong, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Sleuth: Identifying Ponzi Contracts via Large Language Models
by: Wu, Cong, et al.
Published: (2024)
by: Wu, Cong, et al.
Published: (2024)
RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection
by: Wang, Ziwei, et al.
Published: (2026)
by: Wang, Ziwei, et al.
Published: (2026)
Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation
by: Liang, Ruichao, et al.
Published: (2025)
by: Liang, Ruichao, et al.
Published: (2025)
MagLive: Robust Voice Liveness Detection on Smartphones Using Magnetic Pattern Changes
by: Sun, Xiping, et al.
Published: (2024)
by: Sun, Xiping, et al.
Published: (2024)
SCR-Auth: Secure Call Receiver Authentication on Smartphones Using Outer Ear Echoes
by: Sun, Xiping, et al.
Published: (2024)
by: Sun, Xiping, et al.
Published: (2024)
EvoPoC: Automated Exploit Synthesis for DeFi Smart Contracts via Hierarchical Knowledge Graphs
by: Liang, Ruichao, et al.
Published: (2026)
by: Liang, Ruichao, et al.
Published: (2026)
WAFBOOSTER: Automatic Boosting of WAF Security Against Mutated Malicious Payloads
by: Wu, Cong, et al.
Published: (2025)
by: Wu, Cong, et al.
Published: (2025)
Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
by: Wang, Zhilong, et al.
Published: (2024)
by: Wang, Zhilong, et al.
Published: (2024)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
Benchmarking ZK-Friendly Hash Functions and SNARK Proving Systems for EVM-compatible Blockchains
by: Guo, Hanze, et al.
Published: (2024)
by: Guo, Hanze, et al.
Published: (2024)
Disassembling Obfuscated Executables with LLM
by: Rong, Huanyao, et al.
Published: (2024)
by: Rong, Huanyao, et al.
Published: (2024)
ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-based Systems
by: Zhou, Mingyi, et al.
Published: (2023)
by: Zhou, Mingyi, et al.
Published: (2023)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
by: Wang, Zhilong, et al.
Published: (2024)
by: Wang, Zhilong, et al.
Published: (2024)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating Intent
by: Shang, Shang, et al.
Published: (2024)
by: Shang, Shang, et al.
Published: (2024)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
by: Wang, Zeng, et al.
Published: (2026)
by: Wang, Zeng, et al.
Published: (2026)
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
by: Chen, Luoyu, et al.
Published: (2026)
by: Chen, Luoyu, et al.
Published: (2026)
Obfuscating IoT Device Scanning Activity via Adversarial Example Generation
by: Li, Haocong, et al.
Published: (2024)
by: Li, Haocong, et al.
Published: (2024)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
Universal Jailbreak Suffixes Are Strong Attention Hijackers
by: Ben-Tov, Matan, et al.
Published: (2025)
by: Ben-Tov, Matan, et al.
Published: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
by: Gu, Haoran, et al.
Published: (2026)
by: Gu, Haoran, et al.
Published: (2026)
Sealing the Audit-Runtime Gap for LLM Skills
by: Shen, Tingda, et al.
Published: (2026)
by: Shen, Tingda, et al.
Published: (2026)
SecureNT: Smart Topology Obfuscation for Privacy-Aware Network Monitoring
by: Du, Chengze, et al.
Published: (2024)
by: Du, Chengze, et al.
Published: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
CLOAQ: Combined Logic and Angle Obfuscation for Quantum Circuits
by: Langford, Vincent, et al.
Published: (2026)
by: Langford, Vincent, et al.
Published: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Understanding and Characterizing Obfuscated Funds Transfers in Ethereum Smart Contracts
by: Sheng, Zhang, et al.
Published: (2025)
by: Sheng, Zhang, et al.
Published: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation
by: Lazo, Luis, et al.
Published: (2026)
by: Lazo, Luis, et al.
Published: (2026)
ORCAS: Obfuscation-Resilient Binary Code Similarity Analysis using Dominance Enhanced Semantic Graph
by: Wang, Yufeng, et al.
Published: (2025)
by: Wang, Yufeng, et al.
Published: (2025)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
by: Huang, Xun, et al.
Published: (2026)
by: Huang, Xun, et al.
Published: (2026)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Pushan: Trace-Free Deobfuscation of Virtualization-Obfuscated Binaries
by: Sudhir, Ashwin, et al.
Published: (2026)
by: Sudhir, Ashwin, et al.
Published: (2026)
GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
by: Zhang, Zaixi, et al.
Published: (2025)
by: Zhang, Zaixi, et al.
Published: (2025)
Red Teaming Methodology for Design Obfuscation
by: Liu, Yuntao, et al.
Published: (2025)
by: Liu, Yuntao, et al.
Published: (2025)
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
by: Yan, Bo, et al.
Published: (2026)
by: Yan, Bo, et al.
Published: (2026)
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
by: Yildiz, Alperen, et al.
Published: (2025)
by: Yildiz, Alperen, et al.
Published: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
Sok: Comprehensive Security Overview, Challenges, and Future Directions of Voice-Controlled Systems
by: Xu, Haozhe, et al.
Published: (2024)
by: Xu, Haozhe, et al.
Published: (2024)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
by: Zhang, Junke, et al.
Published: (2026)
by: Zhang, Junke, et al.
Published: (2026)
Similar Items
-
Semantic Sleuth: Identifying Ponzi Contracts via Large Language Models
by: Wu, Cong, et al.
Published: (2024) -
RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection
by: Wang, Ziwei, et al.
Published: (2026) -
Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation
by: Liang, Ruichao, et al.
Published: (2025) -
MagLive: Robust Voice Liveness Detection on Smartphones Using Magnetic Pattern Changes
by: Sun, Xiping, et al.
Published: (2024) -
SCR-Auth: Secure Call Receiver Authentication on Smartphones Using Outer Ear Echoes
by: Sun, Xiping, et al.
Published: (2024)