AJAR: Adaptive Jailbreak Architecture for Red-teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Dou, Yipu, Yang, Wang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
by: Xu, Ning, et al.
Published: (2025)
by: Xu, Ning, et al.
Published: (2025)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
by: Yang, Guangyu, et al.
Published: (2025)
by: Yang, Guangyu, et al.
Published: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
by: Pathade, Chetan
Published: (2025)
by: Pathade, Chetan
Published: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
by: Ramesh, Govind, et al.
Published: (2024)
by: Ramesh, Govind, et al.
Published: (2024)
Defending against Jailbreak through Early Exit Generation of Large Language Models
by: Zhao, Chongwen, et al.
Published: (2024)
by: Zhao, Chongwen, et al.
Published: (2024)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
by: Du, Yuhao, et al.
Published: (2024)
by: Du, Yuhao, et al.
Published: (2024)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
by: Zhou, Yukai, et al.
Published: (2025)
by: Zhou, Yukai, et al.
Published: (2025)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
by: Li, Jindong, et al.
Published: (2026)
by: Li, Jindong, et al.
Published: (2026)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
by: Liu, Yanjiang, et al.
Published: (2025)
by: Liu, Yanjiang, et al.
Published: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
by: Yadav, Advait, et al.
Published: (2025)
by: Yadav, Advait, et al.
Published: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
by: Lv, Huijie, et al.
Published: (2024)
by: Lv, Huijie, et al.
Published: (2024)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
by: Zhang, Chiyu, et al.
Published: (2026)
by: Zhang, Chiyu, et al.
Published: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
by: Gu, Haoran, et al.
Published: (2026)
by: Gu, Haoran, et al.
Published: (2026)
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
by: Tshimula, Jean Marie, et al.
Published: (2024)
by: Tshimula, Jean Marie, et al.
Published: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025)
by: Wu, Yu-Hang, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
Similar Items
-
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
by: Dou, Yipu, et al.
Published: (2026) -
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025) -
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024) -
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
by: Xu, Ning, et al.
Published: (2025) -
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
by: Béjar, Mario Rodríguez, et al.
Published: (2026)