DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xinzhe, Xiu, Kedong, Zheng, Tianhang, Zeng, Churui, Ni, Wangze, Qin, Zhan, Ren, Kui, Chen, Chun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025)
by: Xiu, Kedong, et al.
Published: (2025)
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
by: Zeng, Churui, et al.
Published: (2026)
by: Zeng, Churui, et al.
Published: (2026)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
by: Qi, Weiwei, et al.
Published: (2025)
by: Qi, Weiwei, et al.
Published: (2025)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
by: Hu, Yuke, et al.
Published: (2023)
by: Hu, Yuke, et al.
Published: (2023)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
by: Qi, Weiwei, et al.
Published: (2026)
by: Qi, Weiwei, et al.
Published: (2026)
Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
by: Ma, Binhao, et al.
Published: (2024)
by: Ma, Binhao, et al.
Published: (2024)
DV-FSR: A Dual-View Target Attack Framework for Federated Sequential Recommendation
by: Qin, Qitao, et al.
Published: (2024)
by: Qin, Qitao, et al.
Published: (2024)
Combating Concept Drift with Explanatory Detection and Adaptation for Android Malware Classification
by: He, Yiling, et al.
Published: (2024)
by: He, Yiling, et al.
Published: (2024)
Membership Inference Attacks Against Vision-Language Models
by: Hu, Yuke, et al.
Published: (2025)
by: Hu, Yuke, et al.
Published: (2025)
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
The IoT Breaches your Household Again
by: Bonaventura, Davide, et al.
Published: (2024)
by: Bonaventura, Davide, et al.
Published: (2024)
DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern
by: Pang, Xiaoyi, et al.
Published: (2026)
by: Pang, Xiaoyi, et al.
Published: (2026)
Incentivizing Collaboration for Detection of Credential Database Breaches
by: Nanda, Mridu, et al.
Published: (2025)
by: Nanda, Mridu, et al.
Published: (2025)
SWAT: A System-Wide Approach to Tunable Leakage Mitigation in Encrypted Data Stores
by: Zheng, Leqian, et al.
Published: (2023)
by: Zheng, Leqian, et al.
Published: (2023)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
by: Xu, Ning, et al.
Published: (2025)
by: Xu, Ning, et al.
Published: (2025)
Jailbreak Attack Initializations as Extractors of Compliance Directions
by: Levi, Amit, et al.
Published: (2025)
by: Levi, Amit, et al.
Published: (2025)
FDINet: Protecting against DNN Model Extraction via Feature Distortion Index
by: Yao, Hongwei, et al.
Published: (2023)
by: Yao, Hongwei, et al.
Published: (2023)
A Certified Robust Watermark For Large Language Models
by: Feng, Xianheng, et al.
Published: (2024)
by: Feng, Xianheng, et al.
Published: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
Explainer-guided Targeted Adversarial Attacks against Binary Code Similarity Detection Models
by: Chen, Mingjie, et al.
Published: (2025)
by: Chen, Mingjie, et al.
Published: (2025)
Lightweight and Breach-Resilient Authenticated Encryption Framework for Internet of Things
by: Nouma, Saif E., et al.
Published: (2025)
by: Nouma, Saif E., et al.
Published: (2025)
CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
by: Hu, Jiaming, et al.
Published: (2025)
by: Hu, Jiaming, et al.
Published: (2025)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
by: Zeng, Xiyu, et al.
Published: (2025)
by: Zeng, Xiyu, et al.
Published: (2025)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
by: Zhang, Wenhui, et al.
Published: (2025)
by: Zhang, Wenhui, et al.
Published: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
EVA: Editing for Versatile Alignment against Jailbreaks
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
by: He, Yiling, et al.
Published: (2023)
by: He, Yiling, et al.
Published: (2023)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
Accelerating Incident Response: A Hybrid Approach for Data Breach Reporting
by: Arrus, Aurora, et al.
Published: (2026)
by: Arrus, Aurora, et al.
Published: (2026)
Enterprise Security Incident Analysis and Countermeasures Based on the T-Mobile Data Breach
by: Cui, Zhuohan, et al.
Published: (2025)
by: Cui, Zhuohan, et al.
Published: (2025)
From Balance to Breach: Cyber Threats to Battery Energy Storage Systems
by: Öhrström, Frans, et al.
Published: (2025)
by: Öhrström, Frans, et al.
Published: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
by: Huang, Xun, et al.
Published: (2026)
by: Huang, Xun, et al.
Published: (2026)
Crisis Communication in the Face of Data Breaches
by: Ruohonen, Jukka, et al.
Published: (2024)
by: Ruohonen, Jukka, et al.
Published: (2024)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
by: Chen, Tailun, et al.
Published: (2025)
by: Chen, Tailun, et al.
Published: (2025)
Breaking the Vault: A Case Study of the 2022 LastPass Data Breach
by: Gentles, Jessica, et al.
Published: (2025)
by: Gentles, Jessica, et al.
Published: (2025)
A Critical Analysis of the Medibank Health Data Breach and Differential Privacy Solutions
by: Cui, Zhuohan, et al.
Published: (2026)
by: Cui, Zhuohan, et al.
Published: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Similar Items
-
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025) -
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
by: Zeng, Churui, et al.
Published: (2026) -
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025) -
MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
by: Qi, Weiwei, et al.
Published: (2025) -
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)