MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Weiwei, Shao, Shuo, Gu, Wei, Zheng, Tianhang, Zhao, Puning, Qin, Zhan, Ren, Kui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
by: Zeng, Churui, et al.
Published: (2026)
by: Zeng, Churui, et al.
Published: (2026)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
by: Qi, Weiwei, et al.
Published: (2026)
by: Qi, Weiwei, et al.
Published: (2026)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025)
by: Xiu, Kedong, et al.
Published: (2025)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
by: Ma, Binhao, et al.
Published: (2024)
by: Ma, Binhao, et al.
Published: (2024)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
FedTracker: Furnishing Ownership Verification and Traceability for Federated Learning Model
by: Shao, Shuo, et al.
Published: (2022)
by: Shao, Shuo, et al.
Published: (2022)
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
by: Shao, Shuo, et al.
Published: (2024)
by: Shao, Shuo, et al.
Published: (2024)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
by: Chen, Tailun, et al.
Published: (2025)
by: Chen, Tailun, et al.
Published: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
by: Zhao, Tianhang, et al.
Published: (2025)
by: Zhao, Tianhang, et al.
Published: (2025)
SWAT: A System-Wide Approach to Tunable Leakage Mitigation in Encrypted Data Stores
by: Zheng, Leqian, et al.
Published: (2023)
by: Zheng, Leqian, et al.
Published: (2023)
FDINet: Protecting against DNN Model Extraction via Feature Distortion Index
by: Yao, Hongwei, et al.
Published: (2023)
by: Yao, Hongwei, et al.
Published: (2023)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
by: Wang, Xinkai, et al.
Published: (2025)
by: Wang, Xinkai, et al.
Published: (2025)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
by: Nian, Yi, et al.
Published: (2025)
by: Nian, Yi, et al.
Published: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
SoK: On the Role and Future of AIGC Watermarking in the Era of Gen-AI
by: Ren, Kui, et al.
Published: (2024)
by: Ren, Kui, et al.
Published: (2024)
Combating Concept Drift with Explanatory Detection and Adaptation for Android Malware Classification
by: He, Yiling, et al.
Published: (2024)
by: He, Yiling, et al.
Published: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
by: Zhang, Wenhui, et al.
Published: (2025)
by: Zhang, Wenhui, et al.
Published: (2025)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
EVA: Editing for Versatile Alignment against Jailbreaks
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
by: He, Yiling, et al.
Published: (2023)
by: He, Yiling, et al.
Published: (2023)
Membership Inference Attacks Against Vision-Language Models
by: Hu, Yuke, et al.
Published: (2025)
by: Hu, Yuke, et al.
Published: (2025)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
by: Jeong, Joonhyun, et al.
Published: (2025)
by: Jeong, Joonhyun, et al.
Published: (2025)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
FedReview: A Review Mechanism for Rejecting Poisoned Updates in Federated Learning
by: Zheng, Tianhang, et al.
Published: (2024)
by: Zheng, Tianhang, et al.
Published: (2024)
Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
Defense against Poisoning Attacks under Shuffle-DP
by: Wang, Siyi, et al.
Published: (2026)
by: Wang, Siyi, et al.
Published: (2026)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
by: Zeng, Xiyu, et al.
Published: (2025)
by: Zeng, Xiyu, et al.
Published: (2025)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
by: Cui, Tiehan, et al.
Published: (2025)
by: Cui, Tiehan, et al.
Published: (2025)
Similar Items
-
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
by: Zeng, Churui, et al.
Published: (2026) -
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
by: Qi, Weiwei, et al.
Published: (2026) -
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025) -
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025) -
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
by: Huang, Xinzhe, et al.
Published: (2025)